Posted by aagha 9 hours ago
I built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
Definitely! I myself have been having a lot of fun[0] recently using LLMs to enrich one of my favorite beginner Latin readers[1] with audio forced-alignment, POS tagging, morphology, definitions, and other niceties that learners might appreciate.
[0]: https://hercules.hookbangsplat.com
[1]: https://archive.org/details/p1fablesoforbili00godl/mode/2up
Also most of the word popups use different spellings, like "iam" becomes "jam".
For example:
https://en.wikipedia.org/wiki/File:Trajan_inscription_duoton... (Latin as typically carved in stone)
https://commons.wikimedia.org/wiki/File:Herculanean_Rolls_-_... (Greek as typically written on papyrus)
(Source: not an expert in any way)
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
How did you get here?
I have many interests but botany became a big one, and the scientific naming scheme is mostly Latin and Greek, which leads to some interesting word history. "Sativa" means anything "cultivated" (example, "Avena sativa" is "cultivated oats" and "Crocus sativa" is "cultivated saffron"). "Nemo" means "no one" (so the "Search for Nemo" is the search for no one, haha), the root word for "service" means "slave". There are many more interesting associations.
I find the deepest and most awe-inspiring truths within physics (out of which biology and chemistry are arguably sort of abstractions in terms of order of complexity). Behind this is a long awkward history of curious people fumbling for answers in the dark. I am an odd nut, but it is connecting in a comforting way the way I can see others have asked the same burning questions of the "how" and the "why" of things, and the myriad subjects it has lead them to.
As John Muir once said: "When we try to pick out anything by itself, we find it hitched to everything else in the universe."
[0]: https://en.wikipedia.org/wiki/Lingua_Latina_per_se_illustrat...
A lot of role playing games, especially from Square Enix,leverage a lot of Latin and classical education.
The thing that strikes me from ancient texts is not the obvious differences from our day, but the similarities.
May I point out that in "Crudam si edes, in acetum intinguito", that "edes" is more likely to be the future of edere/esse "to eat"? (Just guessing by context.)
If these 2 are fixed I’d use it daily
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
ΕΝ ΑΡΧΗ ἦν ὁ λόγος , καὶ ὁ λόγος ἦν πρὸς τὸν θεόν , καὶ θεὸς ἦν ὁ λόγος .
A propos of breathing marks, iota subscripts, and the three different accent marks of Classical Greek, when I learned it at school we had to remember the breathing marks and iota subscripts, and would lose marks if we omitted them, but we didn't need to learn the accents. Modern Greek now has only (acute) accents, which you need to know to stress the correct vowels, exactly where the accents were in the equivalent Classical Greek words.
What should actually be done, and what OP should do, is take the Perseus website and make it so it doesn't 503 all the time (it is incredibly unreliable).
Then, FOIA the State of California for the pay-to-play data which the UC Irvine-based TLG hoarders are withholding from the public (it is a publicly funded project...) so that the corpus can be meaningfully extended and built upon.
The 1990's tier html vibe of Perseus is to its great advantage.
For tbos version, I gather that a bilingual presentation would be more than necessary: keep Ancient Greek text on a side and display a scholar translation into a selection of switcheable languages
Ancient Library is some vibecoded alternative that the LLM translation it uses will be incorrect. Color me unimpressed.
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.