Case study · 2024–Present · Sole developer
Grammario
Grammar you can see, not just memorize. Universal Dependencies for the hard truth. AI for the teaching.
L'ho fatta parlare in italiano.
Click a word
Past participle
Agrees in gender with the clitic. The structural head of the clause.
Language learning is a core hobby. Across Duolingo, LingQ, textbooks, and tutors, grammar was always explained as rules to memorize, not structures to see. When I analyzed a sentence in my head, I was drawing relationships. No tool reflected that.
That Italian sentence is the drawing I wanted inspectable: pronoun, auxiliary, agreeing past participle, infinitive, preposition, noun.
Structural-First Analysis
Earlier versions asked an LLM to identify grammar. Output was fluent and hallucinatory. I rebuilt the engine in December 2025.
01 · Analyst
spaCy parses via Universal Dependencies, with Stanza as an automatic fallback. Lemmatization, POS tags, and dependency arcs are extracted deterministically. No model hallucination at this layer.
02 · Strategist
Language-specific post-processing. Turkish, German, Russian, Italian, and Spanish each get rules for how they actually build meaning. A one-size-fits-all engine is a mistake.
03 · Tutor
Only after structure is known does the AI explain it in natural language, teaching from a structure that is already on the page.
Six languages shipped
Each language gets its own Strategist pass, because agglutinative and fusional languages don't break down the same way.
- ITItalianAgreement clusters, fusional morphology
- DEGermanCase governance, verb-bracket structures
- RURussianSix-case system, aspect pairs
- TRTurkishAgglutinative X-Ray, suffix decomposition
- ESSpanishAgreement clusters, ser/estar distinction
- JAJapaneseVerb and adjective conjugation, honorific register
Built out from there
- Interactive SVG dependency tree. Click a word for POS, lemma, case, tense, relation.
- Rule-based grammar error detection and CEFR difficulty scoring (A1–C2), both computed from the parsed structure.
- Sentence similarity via pgvector and sentence-transformers against the user's own history.
- Learn section: CEFR-organized grammar curriculum, A1–C2.
- A dual spaced-repetition system for vocabulary and grammar concepts, wrapped in streaks and achievements.
- Teacher suite: classes, live real-time quizzes, AI-graded writing prompts, a gradebook, and a per-student grammar readiness heat map.
- Sentence Remix: past tense, negative, plural, formal, passive, each re-parsed and diff-highlighted against the original.
- Paragraph mode, error-trend analytics, and AI study plans tracked against live mastery data.
Making the AI fast and dependable
Eight generative services run through one Python LLM service with structured JSON outputs, OpenRouter as the primary provider and OpenAI as fallback. The engineering is in everything around the model call.
9s → 4s
full analysis
300–500ms
tree on screen
3 tiers
of caching
Parallel inference
Sequential parse plus LLM took nine seconds. Parsing, the LLM explanation, and the embedding now run concurrently with asyncio.gather over executor futures, about four seconds wall-clock. If the LLM fails, the tree still returns.
Two-phase streaming
A separate NLP-only endpoint puts the dependency tree on screen in 300–500ms while the pedagogy panel fills in behind a skeleton.
Cache before you call
Analyses cache in Redis for 24 hours under SHA-256 keys. Conjugation paradigms are generated once and persisted in Postgres and Redis, so exploring a common verb assembles sentences without calling a model at all.
Rules where rules win
CEFR difficulty is scored from engineered features like tree depth, subordination, and morphological complexity. Quizzes are generated by rules behind a quality gate, and the LLM is only a fallback. Only raw sentence text is sent to providers, never account data.