11 ms·
First word discovered in unopened Herculaneum scroll by CS student
- ejlxsh 3y agoNot bad seeing as he solved this while working as an intern at SpaceX too!
- sillysaurusx 3y agoSee also Nat’s twitter announcement: https://twitter.com/natfriedman/status/1712470683207532906 https://twitter.com/natfriedman/status/1712470683207532906 $700k is a life changing amount of money. I admit, it’s tempting to drop everything and go devote myself like a monk to the pursuit of ancient enlightenment via modern ML. I wonder where we’d start… It’s also funny that the scroll might just be a laundry list.
- chakintosh 3y agoOr a customer complaint: https://www.thearchaeologist.org/blog/complaint-tablet-to-ea-nasir-the-oldest-recorded-customer-complaint https://www.thearchaeologist.org/blog/complaint-tablet-to-ea...
- Wojtkie 3y agoWhat I love about the Ea-Nasir story is the tablet was found in a pile of other tablets, suggesting that Ea-Nasir saved them. Why? Who knows, maybe he found them funny.
- vimax 3y agoI heard somewhere it was common practice to reuse tablets. It was easier to scrape the surface clean than to make a new tablet. You'd save any tablets you have, and might wait until you need it to scrape it clean. In Mesopotamia there was a period where it was fashionable to use a more rare softer red clay on top of the white clay. Your stylus would cut through the top layer leaving nice white letters on a red background. It made it easier to scrape clean and reuse, but much less durable over time.
- jdminhbg 3y agoYes, the clay tablets were used over and over. The ones that are preserved have what was written on them when they were fired, accidentally, by being in a building that was destroyed by fire.
- cheese_van 3y agoYeah, and a large number of clay tablets found were training tablets. At any given time, there may have been more students working on properly learning the mechanics of writing on tablets then tablets in circulation so a surprising number of what is found is, basically homework, or apprentice practice.
- daveevad 3y agoThat's fascinating to this lay clay tablet thinker. It had never occurred to me that clay tablets were reused. In fact, it nearly fundamentally changes my estimation of them as a long-term durable storage medium.
- PeterisP 3y agoHardening clay is relatively simple (just as any other pottery) but optional - you can choose to keep the clay unbaked if you want to reuse it (it would dry out but can be made wet again), but you can also make them permanent.
- user3939382 3y ago> You'd save any tablets you have, and might wait until you need it to scrape it clean. The 3.5” floppies of yore.
- spiritplumber 3y agoThe first meme dump.
- wut-wut 3y agoWell done.
- dmix 3y ago> What do you take me for that you treat me with such contempt? … Apparently "What do you take me for" is an extremely old phrase. Funny how things stick around. I wonder if that's a result of translation though.
- omoikane 3y agoIf you like the Ea-Nasir story, there is a whole subreddit dedicated to it: https://www.reddit.com/r/ReallyShittyCopper/ https://www.reddit.com/r/ReallyShittyCopper/ Also: https://xkcd.com/2758/ https://xkcd.com/2758/
- 0xf00ff00f 3y agoA laundry list with something purple...
- empath-nirvana 3y agoit might cost more than $700k in compute.
- latchkey 3y agoIt certainly did. https://news.ycombinator.com/item?id=36312385 https://news.ycombinator.com/item?id=36312385
- cosmojg 3y agoWhere do they say that the winners used that cluster?
- latchkey 3y agoIt is an assumption based on the fact that the codebase uses cuda and the main backer of the project owns the cluster.
- nyssos 3y ago> It is an assumption Then don't say "certainly"
- latchkey 3y agoWhy not? If they had bought the compute themselves, it might cost more than $700k.
- BHSPitMonkey 3y agoI am absolutely certain that the compute might have cost more than 100 trillion dollars.
- mr_toad 3y agoMany amateurs compete in Kaggle, most of them use whatever hardware they have to hand, and a lot of them will use CUDA directly or indirectly.
- terhechte 3y agoThis is one of more than 600 scrolls that could be read afterwards if the method becomes scalable. What's more: "excavations were never completed, and many historians believe that thousands more scrolls remain underground." [0] [0]: https://scrollprize.org https://scrollprize.org
- ssnistfajen 3y agoBeing able to virtually unroll the scrolls within reasonable time/effort/cost will hopefully encourage the Italian government to approve further excavations of the villa for more papyri. If Vesuvius erupts again we don't know how much of the present excavation will survive.
- radarsat1 3y ago> if the method becomes scalable the machine learning stuff is cool, but it's important not to discount the apparently pretty manual labour still involved in the virtual unwrapping: > Early in the summer, a small team of annotators (the “segmentation team”) joined our effort. They began mapping the 3D structure of the scroll using tools initially built by EduceLab and improved by our community. By July we had segmented and “virtually flattened” hundreds of cm2 of papyrus. So, it sounds like it was about a month or two of work, for a single scroll. Although, it probably could be partially or fully automated too, with some effort. Already they developed some tools to help, and I guess it's the kind of task that gets easier after you do it the first time.
- versteegen 3y agoActually it's much worse than that. Only a very small fraction of the first of the two scanned scrolls has been segmented/unwrapped after 5 months, and it's the easiest parts that are done -- about 1000cm^2 across something like 100 layers of papyrus 10cm wide. Only 50cm^2 of scroll 2 is done. Where the sheets are right against each other is much harder.
- gojomo 3y agoBut at the same time: scanning tech & software automation just keep getting better, including via spillovers from other unrelated projects. The ability of an ML system to learn to mimic what the manual "virtual unrolling" process is doing, from a small number of examples-to-follow, is growing. Each bit of success, once confirmed by other experts or correlation with other texts, improves the training data. Eventually a fully-software pushbutton pipeline of "raw imaging to likely texts" should be possible. And if, say, some of the scrolls are sufficiently 'read' nondestructively to embolden teams to risk destructive techniques – such as incremental ablation while reading the exact chemicals at every coordinate – even higher-resolution data could become available.
- jdminhbg 3y ago> It’s also funny that the scroll might just be a laundry list. Most likely not, I believe they're starting with scrolls that were readable on the outside, which we know are minor works of Greek stoic philosophy. Also a laundry list would be written on a reusable wax tablet, rather than costly papyrus.
- michael_nielsen 3y agoIt's likely somehow a reference to the Emperor. Purple cloth was extremely rare and expensive, and it was the colour worn by the Emperors. Indeed, it eventually became a capital crime for people outside the Emperor's family to wear it. I don't know if that was yet true at the time of Vesuvius, although Wikipedia claims Caligula may have had someone killed for wearing purple.
- Arete314159 3y agoThe other word visible is "oino", wine. Wine can be described as purple.
- thaumasiotes 3y agoOnly if you assume that whoever scribed the scroll wasn't too concerned about what order the letters in a word should be written in. The image is annotated OIWN and the article tentatively identifies the word as OMOIWN, meaning "similar".
- londons_explore 3y agowine stains are decidedly more purple than the wine they came from too.
- OfSanguineFire 3y agoWhile modern people make that connection, that is culturally dependent. The color terms available to speakers of a language, and what objects those terms can be associated with, change over time. In the case of the Greek word for "purple", it was connected to a dye and therefore used for clothing, but one shouldn't expect it to be used for wine.
- fsckboy 3y agohttps://en.wikipedia.org/wiki/Wine-dark_sea_(Homer) https://en.wikipedia.org/wiki/Wine-dark_sea_(Homer) "Wine-dark sea is a traditional English translation of oînops póntos (οἶνοψ πόντος, IPA: /ôi̯.nops pón.tos/), from oînos (οἶνος, "wine") + óps (ὄψ, "eye; face"), a Homeric epithet. A literal translation is "wine-face sea" (wine-faced, wine-eyed). It is attested five times in the Iliad and twelve times in the Odyssey[1] often to describe rough, stormy seas. The only other use of oînops in the works of Homer is for oxen, for which is it used once in the Iliad and once in the Odyssey, where it describes a reddish colour. The phrase has become a common example when talking about the use of colour in ancient Greek texts."
- dataflow 3y ago> $700k is a life changing amount of money Probably ~half of that will go to taxes?
- thrway63245 3y agoNot sure why this is downvoted. Yes, in California half will go to taxes and the rest is enough for a downpayment on a shack. Hardly life changing.
- lazyasciiart 3y agoPretty damn life changing, actually, for the vast majority of even gasp Californians.
- HeyLaughingBoy 3y agoMust be nice to be a billionaire. I'm a fairly well paid engineer and I would certainly consider a $700k gift to be life-changing.
- hotnfresh 3y agoI make good, but not FAANG, money, and $350k after-tax, all at once, would still put me most of a decade ahead, financially, and give me some much better options for dealing with a bunch of stuff. Yeah, life changing.
- dataflow 3y ago> would still put me most of a decade ahead, financially I think that might be the kind of the crux of it. Saving this money to get a decade "ahead" might help your retirement life (assuming things don't collapse before then), but not your current life. Whether that's considered "life-changing" might be in the eye of the beholder.
- lazide 3y agoNo one can not consider that life changing in that context?
- jedberg 3y ago> It’s also funny that the scroll might just be a laundry list. Even if it were, a laundry list from 2000 years ago would be a fascinating read.
- KingLancelot 3y ago[dead]
- dclowd9901 3y agoI don’t think people in those days wrote bullshit. You had to know how to write, and if you did, no one was entreating you to write grocery lists.
- gambiting 3y agoThey absolutely did, we have plenty of proof of ancient romans and other cultures writing down jokes, little squabbles, "John was here" on walls etc, one of the oldest known pieces of writing is literally one merchant complaining to another merchant about some marble that wasn't as ordered.
- 1-more 3y agoon walls and tablets, but on papyrus? Papyrus was expensive.
- jakderrida 3y ago>I admit, it’s tempting to drop everything and go devote myself like a monk to the pursuit of ancient enlightenment via modern ML. I wonder where we’d start… I think you'd be shocked how well LLMs translate cuneiform in the CDLI notation. What's hilarious is my first attempt included examples in-context and Claude prefaced the translation by stating that there's nothing in my example translations about "bulls", "horns" or "grabbing" and that it will ignore that translation. I looked it up word-by-word and realized Claude was right. Blew me away. Yet Assyriology subreddits were as excited about my findings as lawyer subreddits are about LLMs. Not sure why, either. Just a bunch of, "So what? Does that mean it's useful?".
- davidw 3y agoThat's extremely cool. I wonder what we'll learn. As an aside, the "Professor Seales and team scanning at the particle accelerator" photo looks like it came from a TV show. "If we keep telling the computer 'enhance', we'll be able to read it".
- adamlgerber 3y agoi love this project. i feel like this is going to be a great source of interest and value over the next few years (and potentially immesurable value over longer time frames).
- kelsey9876543 3y agoI recently saw a wonderful youtube video on this: https://www.youtube.com/watch?v=Z_L1oN8y7Bs https://www.youtube.com/watch?v=Z_L1oN8y7Bs Title: Herculaneum scrolls: A 20-year journey to read the unreadable it goes a little bit into the technology of how this was done, deep learning finally cracked the code. They had the scans for a decade but it took ML training to be able to identify which parts were paper and which parts were the ink on top. This had been done on a different set of scrolls with easier to read higher contrasting materials like the video says, 20 years ago. Deep learning is cracking the code for these datasets we had previously thought were impossible to algorithmically solve.
- nulbyte 3y agoThank you for sharing. It's a month old, but even so, I just saw a pinned comment ppsted an hour ago about an announcement coming later today.
- versteegen 3y agoCan't speak for the video, but this is a bit misleading actually. What cracked this was actually visual inspection looking for patterns which could then be used as better training data, which so far apparently hasn't found very many letters that were too hard to see. Read the OP describing the iterative process of hand-annotation guided by output of a model, then retraining the model with the additional data, it's a fascinating technique! Simply using deep learning on the initially available ground truths without knowing what features the models should be looking for actually pretty much didn't work! Also, so far the process of virtually unrolling the scrolls is mostly manual and extremely labour intensive.
- kelsey9876543 3y agoThank you for adding the deeper insight! The competition and the methods used are very fascinating indeed.
- Etheryte 3y agoThis is highly misleading. Deep learning was not what did the discovery, the find was handmade. They're trying to make a deep learning model do what was done by hand here, but so far they haven't had success in it finding actual letters.
- tclancy 3y agoSomewhat off-topic but if you clicked in here, you might be interested in this book: "The Riddle of the Labyrinth: The Quest to Crack an Ancient Code".
- jdminhbg 3y agoThis is the 21st-century equivalent of living through the opening of Tut's tomb. Incredible to think there's a very real chance that in the medium-term future you might be able to buy a copy of a newly-translated work on Amazon that hasn't been read for millennia.
- carapace 3y agoWhy the ad for Amazon?
- jdminhbg 3y agoIt's just a reference to making a boring, pervasive part of culture. Please feel free to buy those translations at any book company you feel like.
- carapace 3y agoSorry, I'm just cranky this morning.
- marktani 3y agoI hope it's also true that you're cranky just this morning
- alanbernstein 3y agoSurely they will be public domain by now??
- lexicality 3y agoIt is disgraceful that the ancient Greek authors won't see an obol that these so called "translators" and "historians" make from reselling their work. They should sue! /s
- Ylpertnodi 3y ago
- versteegen 3y agoThe lettering was found by looking for 'crackle' texture on papyrus segments from the CT scans which obviously were in the shape of Greek letters, and annotating those as training data. Unfortunately such crackle texture isn't visible, at least by eye, on most of the papyrus. Probably it's only that visible where the ink was very thick. You can easily see the difference in texture in this electron microscope image [1] (far higher resolution than the CT scans) but especially on the very edge of the inked area (the narrow strip in the left image; I think the whole right image is inked) where the ink was pushed to. I'm surprised the crackle was discovered only after the Kaggle Ink Detection contest. Looking at the CT-scanned fragments with infrared ground truths, which were used in the Kaggle contest, Casey Handmer wrote [2]: > The ongoing apparent failure of deep-learning based ink detection based on the fragments indicated to me that direct inspection of the actual data would be more fruitful, as it has been here. > ... > I found similar “cracked mud” and “flake” textures corresponding to known character ink, but only for perhaps 10% of the known characters. It’s been a long day, I can probably find more on closer inspection, but that does make one wonder about automated ink detection and what that is seeing. These new images are much better than I hoped for, but still only in one small area, so I'm still pessimistic about more than an odd sentence being readable. [1] https://scrollprize.org/img/tutorials/sem.png https://scrollprize.org/img/tutorials/sem.png [2] https://caseyhandmer.wordpress.com/2023/08/05/reading-ancient-scrolls/ https://caseyhandmer.wordpress.com/2023/08/05/reading-ancien...
- tysam_and 3y agoI actually participated in the challenge for a little while and this was the approach I took before I dropped out to do a few other things. However, what I did was a bit different -- instead of looking for a crackle, I surmised that that 'crackling' effect actually is just of course slices of the data over different rifts in the parchment, and that the data of the ink lay on the manifold of that crackling and bending. It would not be as clear to the human eye for all of the letters, I think, as there are many, many, many layers in the scanned image, and you can only start to see a pattern emerge over time as you cycle through the images. I was working on code that minimized an optimization function that was basically the total variance loss if I recall correctly, where it just interpolated each pixel column up and down bilinearly to 'align' the blocks of the image so that the crackle texture was flattened. From there I planned on using a rather optimized convolutional network on the 'flattened' image, which can I think be done rather efficiently as if you look at a cross section of the scroll you can see where it's like a tree in that the pinching and such seems to be somewhat locally consistent, so you might be able to get away with some interpolation. I should probably share the code if this is of interest to anyone, since I'm not pursuing the competition at the moment. Also, this is why I did not buy into 3D convolutions for this, at least. Ink that has been laid and dried should follow a semi-predictable pattern that a 2D convolution can detect, I do not know if a 3D convolution really brings us anything, as the invariances we desire can be structured up front more easily. If there is interest in the code, let me know and I can do a little digging, otherwise, it is a fun challenge, for sure.
- munificent 3y agoI love uses of machine learning like this a thousand times more than generative LLMs spouting probable-sounding nonsense.
- esafak 3y agoIt is amazing what some college student can pull off with today's technology.
- lukeboi 3y agothanks!
- Rallen89 3y ago>Shortly after that, another contestant, Youssef Nader, independently discovered the same word in the same area, with even clearer results — winning the second place prize of $10,000. That's what u get for optimising your code
- hansoolo 3y agoI thought the same. He had the better results, but too late.
- QuercusMax 3y agoOr maybe the winner optimized his code, resulting in faster time to get results. Either one is equally plausible!
- countrymile 3y agocomputed on his laptop whilst he was at a party apparently. Legend!
- lukeboi 3y agoluke here, thanks!
- esafak 3y agoIf quality is a factor, they should have withheld the prize for a reasonable time (a day?) in case someone posts a better result.
- zeteo 3y agoNot really: >Youssef used a model from the Kaggle competition and was inspired by Luke’s results to look in the same area.
- deleted 3y ago[deleted]
- autokad 3y agoimagine the person making this scroll 2,000 years ago wondering 'I wonder if some kid 2000 years in the future is going to win a boat load of money by reading this'
- nataliste 3y agoI wrote this for a different community (filled with semiliterate sophists), but this is absolutely huge and could upend huge swathes of understanding about the last two thousand years. You can avoid the longform essay below if you want. The short of it is there are several potentially common works possibly in the library that could directly prove or disprove what is found in the New Testament and the predicates of Rabbinic Judaism as established at the Council of Jamnia. We could be seeing the beginning of conclusive proof that invalidates the narratives of Christianity, Judaism, and Islam by the end of the year. The Vesuvius Challenge isn't just an interesting contest in the machine learning realm; it's a groundbreaking endeavor that could redefine our understanding of the humanities if successful. The opportunity to digitally unroll and read the Herculaneum Papyri could offer unprecedented insights into ancient civilizations and the total feedstock of civilization today. This is not merely about filling in some historical gaps; it’s about fundamentally altering how we understand antiquity and, by extension, our own intellectual heritage. The loss of the Library of Alexandria has long been considered a "dark age" event for intellectual progress. Now, consider the Herculaneum library—a collection of papyri from a villa once owned by Julius Caesar's father-in-law, carbonized but preserved by the Vesuvius eruption in 79 AD. Hundreds of these scrolls are unreadable because their carbon-based ink blends in with the carbonized papyrus, and thus are invisible to conventional imaging techniques. Yet, these scrolls are quite possibly on the cusp of revelation. Recent developments have introduced machine learning and high-resolution X-ray scans as methods for reading these "unreadable" scrolls. What texts do they contain? Treatises on science and philosophy? The lost books of Livy? The epic cycle? Governmental policies like the Twelve Tables? It’s a tantalizing question because whatever is locked in those scrolls could be an unfiltered look at the Roman Empire—an empire that fundamentally influenced the trajectory of Western culture, religion, governance, and philosophy. Ponder a history of Rome that has not been retouched by myriadic emperors, by Constantine's Christianity, or the interpretive lens of the Roman Catholic Church. Unmediated accounts of Roman society, unaltered by the layers of religious and political power that came later, could rewrite our textbooks and shift the justification of history. It’s not just about enriching our understanding of ancient civilizations; this could be a cornerstone on which to build a fresh philosophical understanding of human society. If the project succeeds, there will be repercussions in the academic realm. The humanities have long struggled to justify their existence in a world that increasingly prizes STEM and lacks any novel sources for the classical world. Suddenly, there could be a concrete, urgent task at hand: to decode, interpret, and integrate an influx of new knowledge. The Vesuvius Challenge could revitalize the field, offering an unforeseen but compelling reason for its study. In essence, it provides a utilitarian justification for the humanities, one that transcends 'cultural enrichment' and enters the realm of 'historical redefinition.' The Vesuvius Challenge could be the hinge upon which history swings, yielding intellectual treasure that could be as groundbreaking as the writings that were lost in Alexandria. For millennia, those scrolls have remained unread. Now, it's a software problem. That's not just a challenge; it’s an imperative. The presence of specific works in the Herculaneum Papyri could dramatically impact our understanding of major historical events. In particular for me, I pray that the biography of Herod the Great by Nicholas of Damascus is discovered intact. While mainstream accounts generally portray the life of Herod within the context of Roman patronage and Judaean politics, uncovering a contemporary account by a close intimate (and used as a primary source by Josephus) would offer fresh, unmediated insights into his rule and its socio-political intricacies. Chronologies of the life of Jesus could be explicitly validated or disproved. The relevance here is far from academic. Consider the following naturalistic hypothesis: that the inception and rise of Christianity was entirely a dynastic struggle within the Hasmonean-Herodian line. What if the tale of Jesus is, in essence, a dramatized, mystified rendition of a 1st-century dynastic conflict, one that was subsequently co-opted and transformed into a religious narrative by an early form of conspiratorial thinking? Something like a 1st-century version of Q-anon, distorting real events to serve an alternative, concealed agenda in the aftermath of the First Jewish-Roman War. Unveiling a document like Nicholas of Damascus' biography could be groundbreaking in testing such a hypothesis. If Herod's life and rule were detailed without the religious overlays that later Christian interpretations bring into the picture, one could make more definitive assertions about the socio-political environment of the time. Furthermore, it could provide concrete evidence to either substantiate or refute theories about Christianity's emergence as a byproduct of a Herodian-Hasmonean power struggle. The fact that such a theory could be tested is significant in its own right. Traditionally, discussions about early Christianity rely heavily on religious texts and subsequent historical accounts, many of which are fraught with dogma and ideological interpretations. A primary source devoid of such influences would be a game-changer, offering a baseline of raw data from which more accurate and reliable hypotheses could be drawn. And it's not limited solely to Christianity. Rabbinic Judaism could have equally monumental implications as a result. The owner of the villa, likely a wealthy Roman, would be unlikely to have had any primary Hebrew texts like the Pentateuch. However, that doesn't rule out the possibility of possessing Greek or Latin works discussing Jewish culture, beliefs, and politics. Given the villa's historical context, it's conceivable that there might be indirect ethnographic accounts from the period surrounding the destruction of Jerusalem in 70 AD but before the Council of Jamnia, traditionally dated around 90 AD, which helped canonize Hebrew scriptures. Why is this important? The Council of Jamnia is often cited as a crucial moment for the development of Rabbinic Judaism. It allegedly led to the fixing of the Hebrew Bible canon and crystallized what would become Talmudic tradition. If documents were to surface that provide a snapshot of Judaic thought and practice just before this council, it could upend millennia of precedent and identity. In a broader context, discovering pre-Jamnia ethnographic sources could significantly change our understanding of how Judaism adapted and evolved in the aftermath of the Second Temple's destruction. This could lead to far-reaching questions. How much of the Talmudic tradition was actually a post-hoc rationalization or systematization of beliefs and practices that were far more fluid before the Council of Jamnia? How much anti-Romanism was pared away to prevent suppression? Moreover, how would such a revelation interact with or even challenge the validity of current Rabbinic and Orthodox Jewish practices? The implications for the Judeo-Christian heritage as a whole are staggering. If both Christianity and Judaism could be traced back explicitly to politically or socially motivated machinations, rather than divinely inspired or time-honored traditions, the entire foundation of Judeo-Christian culture would come into question. In essence, the Vesuvius Challenge has the potential to destabilize two of the world’s major religious traditions at their historical roots. It is difficult to overstate the potential impacts. The Vesuvius Challenge is not just an academic or technological endeavor. Its success could instigate an unparalleled epistemological crisis in religious studies and the humanities. It provides the opportunity to re-examine, with primary sources, the historical foundations of Western religious, cultural, and ultimately political traditions. We're not just potentially rewriting history here; we're reevaluating the very frameworks through which that history has been understood.
- 1vuio0pswjnm7 3y ago"He found a few dozen ink strokes - and some complete letters - that could be labeled and used as training data. Before long, the model was unveiling traces of crackle invisible to his own eye. Soon, these traces began to form letters and hints of actual words." This does not sound like a "Large Language Model (LLM)" or other large set of training data, like the sort hyped by so-called "tech" companies; this sounds relatively small. What am I missing. (Besides brain cells.)
- deleted 3y ago[deleted]
- comex 3y agoIndeed, it’s a machine learning model, but not a large one. Who called it large?
- 1vuio0pswjnm7 3y agoHe made $40,000 without needing a large data set. No proprietary, corporate LLM needed.
- countrymile 3y agocalculated on his laptop whilst at a party
- tedunangst 3y ago> Casey was the first person in 2,000 years to find ink — and a letter — inside an unopened scroll. Amusing that this implies the Vesuviuans had the ability to read unopened scrolls.
- deleted 3y ago[deleted]
- jacobjwebber 3y agoThe word is Ancient Greek for “purple”..it takes a lot of reading to find out!
- f0e4c2f7 3y agoThe first 30 minutes of this interview does a great job of explaining what's going on here. Really interesting stuff. https://youtu.be/qcvMjoJdck4?si=RL1WnAAj4loS1D2s https://youtu.be/qcvMjoJdck4?si=RL1WnAAj4loS1D2s
- mkoubaa 3y ago> Note that texts from this time didn’t use spaces, making it harder to determine word boundaries. Paging Germans
- konstantinua00 3y agoit's not about _combining_ words into one better page Japanese
- BuyMyBitcoins 3y agoI often wonder how different personal computing and programming would be if we kept using Scriptio Continua or the particularly awkward Boustrophedon where every other line is written in the reverse direction. https://en.wikipedia.org/wiki/Scriptio_continua https://en.wikipedia.org/wiki/Scriptio_continua https://en.wikipedia.org/wiki/Boustrophedon https://en.wikipedia.org/wiki/Boustrophedon
- russellbeattie 3y agoImsurethatwedbeabletoreadtextswrittenwithoutspacesorpunctuationitwouldjustbeabitmoredifficultforbeginnerstoparsethetextbutafteryouvegottenenoughvocabularyitdbeprettystraightforwardtoreadthoughspellcheckerswouldprobablyhaveahardertimetheoneinmybrowseriscompletelydfreakingoutoverthistextanotherthingitwouldprobablyhaveaneffedtoniswordchoicesinceyoudhavetomakesurethewordsyouchosetoputtogethercouldntbemisconstruedorambiguous. The idea of reading text backwards and bɘɿoɿɿim ɿɘɈɈɘl ʜɔɒɘ ʜɈiw ,ƨbɿɒwɿoʇ is absolutely nuts though. I'm definitely .ɈɒʜɈ Ɉqobɒ Ɉ'nbib ɘw bɒlϱ
- timschmidt 3y agoIntentionally choosing words which could be misconstrued when written without spaces could make for some hilarious jokes.
- huytersd 3y agoFor the exact opposite of this, I love what Sanskrit/Hindi does. Each complete word has a bar over it.
- numitus 3y agoFrom my perspective it may be just fitting to the answer. We try to find symbols, without confidence, that the paper still contains information, and any text for trainig. With "rights" network you can achieve any possibile result. It is remember me a russian freak-scientist, which try to read words and texts on the detailed sun surface photos.
- dom96 3y agoThis was my first thought. How sure are we that their model isn't just over-fitting and finding letters that aren't really there?
- numitus 3y agoI think they need to make similar scroll, write any text and make the similar damages to check is the approach worked at all
- lolc 3y agoI wondered about that. My understanding is that the models were trained to look for letter shapes, not words. And that the models couldn't produce known words unless they were trained on the language. If it wasn't trained on a substantial text body, a model producing letter sequences that form known words means it found something and didn't hallucinate.
- userbinator 3y agoThis might count as one of the most extreme stories of data recovery I've seen. I wonder if in another 2000 years we'll have a "first file discovered on discarded hard drive platter".
- sexy_seedbox 3y ago"Spot robot dumpster dives in landfill and finds 100,000 bitcoins in discarded hard drive"
- seanthemon 3y ago"We're working hard to figure out what this 'bitcoin' is, exciting times!"
- userbinator 3y agoGiven the relative volumes of data prevalent today, it's more likely that random files on a hard drive, or fragments thereof, will be from porn or some other video.
- wglb 3y agoReminds me faintly of this story https://en.wikipedia.org/wiki/MS_Fnd_in_a_Lbry#:~:text=MS%20Fnd%20in%20a%20Lbry%20(probably%20intended%20to%20be%20understood,Edgar%20Allan%20Poe%20story%20%22MS https://en.wikipedia.org/wiki/MS_Fnd_in_a_Lbry#:~:text=MS%20....
- pests 3y agoQuick aside: I know Chrome has supported the link-to-highlight for years now, but does anyone know where the "#:~:text=" hash format is documented at? Searching for that is really hard.
- thaliaarchi 3y agohttps://developer.mozilla.org/en-US/docs/Web/Text_fragments https://developer.mozilla.org/en-US/docs/Web/Text_fragments
- utopcell 3y agoThis, along with the way the Antikythera mechanism [1] fragments were decoded [2], always brings to mind the ending of "A.I.: Artificial Intelligence", where aliens far into the future managed to recreate part of human civilization from remaining artifacts and a barely working robot child. [1] https://en.wikipedia.org/wiki/Antikythera_mechanism https://en.wikipedia.org/wiki/Antikythera_mechanism [2] https://www.youtube.com/watch?v=6Wp3wL8g2Eg https://www.youtube.com/watch?v=6Wp3wL8g2Eg
- Izkata 3y agoThe design of those "aliens" always had me thinking they were super-advanced robots uncovering the beginning of their civilization.
- utopcell 3y agoQuite possibly. Also, the flash-forward in that movie was after 2,000 years, which would make the connection with the Antikythera mechanism and the Herculaneum scroll even more relevant. We are the "aliens".
- tgsovlerkhgsel 3y agohttps://scrollprize.org/ https://scrollprize.org/ explains the original challenge/issue: other research showed "virtual unwrapping" based on CT scans was possible, but these scrolls had ink not clearly visible on CT/X-ray, so they had to go back to less visible structure (I'm not sure if it's the structure of the paper that changes or actually the ink being visible but less).
- lettergram 3y agoAre they controlling for this? How is the validation being done?
- dexsst 3y agoOn a side note iirc there were also some Dead Sea scrolls that were hard to open, were they able to open and read all of the scrolls?
- tanh 3y agoBased on this, I wonder if the main challenge has already been solved.
- dr_dshiv 3y agoI highly recommend getting into classical literature. It is incredible and beyond conception — I had a light scattering in my education, but only later and recently have I discovered how incredible it can be. Suggested Reading for beginners: * Life of Pythagoras, by Iamblichus * The Golden Ass, by Apuleius of Numenia (specifically, translation by Robert Graves) * Life of Alexander by Plutarch * Education of Cyrus by Xenophon * Parmenides by Plato Also, I have found SHWEP.net to be invaluable for a gentle yet rigorous guide through many classics, though it takes an esoteric bent (which I love)
- bambax 3y agoAnd now the size of the corpus may be about to explode! From the article: > If these words are indeed what we think they are, this papyrus scroll likely contains an entirely new text, unseen by the modern world.
- atomicnature 3y agoLife of Pythagoras was super influential to my life. I was very inspired by the way Pythagoras lived his life - a 100% commitment to finding out what is true, travelling all over the world amassing knowledge/practices/skills through humility, trying so many things out and finally spending so much of his life teaching. Pythagoras was probably the greatest philosopher ever. I also recommend reading about Archytas, one of his spiritual descendants, his thoughts on math, mechanics, music, learning and philosophy are amazing.
- bonoboTP 3y agoBut he got scared of sqrt(2).
- hax0ras 3y agoI guess you could say he was being irrational about it. Da dum.
- dr_dshiv 3y agoArchytas is famous for having designed a steam powered glider. And for having freed Plato from his brief slavery. But he also appears to have written the first treatise on mechanical engineering: https://digitalcommons.unl.edu/cgi/viewcontent.cgi?article=1067&context=classicsfacpub https://digitalcommons.unl.edu/cgi/viewcontent.cgi?article=1...
- Fleusal 3y agoHow can they be sure that the results are actual words from the scroll and not just hallucinations of the neural network? If the scrolls are in such bad condition that the data is almost only noise then what stops their high-tech deep learning model from just.. making it all up?
- lolc 3y agoAs long as the model doesn't know about words, and finds letter sequences that correspond to words, we can conclude that it found actual writing.
- yieldcrv 3y ago> In early August, contestant Casey Handmer, an ex-JPL startup founder and polymath interesting terminology, I've never been given accolades for being a multifaceted human being. I've gotten "generalist" and "after much consideration, we have decided not to proceed with your candidacy "
- precompute 3y agoHilarious. From what I've seen, genuine polymaths shirk away from being identified as one, and if they do decide to get recognition, they're shooed away for the unforgivable sin of Not Being Famous Enough.
- elif 3y agoI wonder how many AI massage iterations this is away from being able to completely copy arbitrary books without opening them, and if this technology will hasten the demise of paper texts. If you can just stack 20 random books and within seconds have them be indexed and searchable digital ones, libraries as we know them will suffer perhaps the final blow in obsolescence.
- croisillon 3y agoOT: never before have i seen a 2 days old story coming back to the front page