LanguaTalk
Every model was a tradeoff. So LanguaTalk built their own.
Don PottingerCo-founder & CTO, LanguaTalk282ms vs 1,943ms at p50
Same qualitygraded by LanguaTalk's own calibrated judge
3x lower costper request
The assumption
LanguaTalk's flagship product, Langua, gives 30,000+ learners a 24/7 AI conversation partner in 20+ languages. Corrections are one of its most-used AI features: the learner writes a sentence in Spanish, and the app writes it back with the mistake struck through and the fix in bold. Many learners run it automatically, so a fresh correction lands after nearly every message. It has to land fast. A correction that shows up after the conversation has moved on feels intrusive rather than helpful.
The feature ran on Claude Sonnet 4, and like most teams shipping AI features, LanguaTalk assumed the next model would be the better model. Then they tried Sonnet 4.5. It behaved differently with their prompts and was worse in the product, so they rolled back to Sonnet 4. When Sonnet 4.6 came out, the team tested it by hand, working through known failure behaviours one at a time, and moved to it.
Manual testing caught real problems, but it did not scale. The team was starting to run evals, and that is when three.dev came in. What they needed next was process and tooling: a way for the team to agree on what "good" looks like, a way to test new models and prompts quickly and reliably, and a full picture of how each candidate performed on quality, latency, and cost.
Using three.dev has been a huge leap for us in providing our users the highest quality while maintaining fast responses.
Don Pottinger, Co-founder & CTO, LanguaTalkAgree on what "good" means
Corrections have no click-through rate. Nothing in the product says a correction was right, and users are not a reliable source of high-volume feedback on quality.
So LanguaTalk turned to an AI Judge: a model whose only job is to grade the feature's outputs, answering one question on every request: "Is this output acceptable to show to a user?"
They had tried LLM-as-a-judge before and were skeptical, with good reason. Without a way to calibrate the judge and measure its alignment, you are tricking yourself into a false sense of confidence.
We tried using LLM judges before, I just don't have a lot of confidence in what it's returning. Without some way to measure and improve alignment, previous evals were hard to trust.
Don Pottinger, Co-founder & CTO, LanguaTalkThat skepticism is exactly what three.dev's calibration process exists to answer. The judge was measured against pass/fail assessments from LanguaTalk's own domain experts and refined until it agreed with them more than 80% of the time. A judge they could trust is what made the experiment results trustworthy.
Try everything
With a judge they could trust, LanguaTalk no longer had to ask whether the newest Sonnet beat the last one. They could ask the real question: what is the best possible setup for this feature?
The bar was never quality alone. A correction has to land while the learner is still reading, so every candidate has to meet the same requirement the product lives under: a median (p50) latency of about one second, at a cost that makes sense for a feature that fires on nearly every message.
three.dev replayed their real historical traffic through 40+ variants spanning Anthropic, OpenAI, Google, and open-weight models: different models, prompts, reasoning levels, thinking on and off, and temperature. Each variant answered the same conversations, so the comparison was apples to apples.
Every single result was a tradeoff.
Inside the latency and cost the product needed, nothing beat their setup on quality. Every model that beat it on quality was too slow or too expensive to ship.
Build the model you want
If no existing model sits where you need it, the remaining move is to make one. LanguaTalk used three.dev to create a small language model (SLM), fine-tuned on their own corrections data and specialized for exactly one job: Spanish corrections, in their format, aligned with their judgment.
The SLM did not land on the frontier. It moved it. Against Sonnet 4.6, the model in production by then, the SLM won on all three axes: better quality (+13%, with the fewest regressions of any variant tested), 282ms at p50, and 56% lower cost per request.
What happens when the next frontier model ships
The obvious worry about a specialized model: it is tuned to today, and the labs ship every few months. So that got measured too. Each notable frontier release goes through the same replay, against the same judge, head-to-head with the SLM.
Measured on 1,464 real conversations, every variant answering the same inputs, graded by LanguaTalk's calibrated judge. Quality and cost are relative to Sonnet 4.6, the model the SLM replaced.
On quality, the SLM, Sonnet 5, and GPT-5.6 Terra are a statistical tie. On speed it is not close: the SLM is 5.3x faster than the Sonnet 4.6 it replaced, 6.9x faster than Sonnet 5, and 3.6x faster than GPT-5.6 Terra, comfortably inside the one-second budget. GPT-5.6 Terra sits right at that line, and every Sonnet configuration is over it. On cost, the SLM is the cheapest option on the board, at roughly a third of Sonnet 5's bill per request.
The reason is structural. A frontier model has to be good at everything. A model trained on your own data only has to be good at one thing.
Sonnet 5 costing 46% more than Sonnet 4.6 was a surprise, because the two share the same list price. What changed is the tokenizer, the dictionary a model uses to chop text into the chunks you are billed for. Sonnet 5 turns the same request into about 1.4x as many chunks, so the identical request costs about 40% more even at identical per-token pricing.
This was an important learning for understanding the real cost per task. Two models can have the exact same cost per token, but what a request actually costs is task and model dependent. The only way to truly know is to run your own requests through the model and measure.
Same score, different failure modes
Even when quality ties, the failure modes underneath differ. three.dev groups every failed correction into named failure modes, so LanguaTalk could see not only how often each model failed, but how.
Sonnet 5 and GPT-5.6 Terra fail the same way. They tend to "fix" Spanish that was already valid, and they break the correction markup more often. The SLM fails the opposite way. It sometimes praises a sentence that still has an error, and it almost never breaks the markup. Same score, different personalities.
That turned the choice into a product decision. Correcting text that was already right erodes a learner's trust faster than missing an error, so LanguaTalk set over-correction as the more severe failure mode and accepted the SLM's under-correction as the failure to live with.
Where this leaves them
LanguaTalk went from manual testing to a repeatable check: any change to any part of their AI setup, whether model, prompt, or reasoning level, is replayed against real traffic and scored on quality, latency, and cost before a learner ever sees it. Integration was a handful of lines of code.
And they own the model running one of their most-used features, at a combination of quality, latency, and price no provider offers off the shelf. No lab's release schedule, deprecation calendar, or price change decides what their product does next.
What to take from it
- "Newer model, better product" is a guess. Overall quality can go up while the failures your users feel most get worse. Knowing which failures matter takes your team's expertise.
- The tradeoff is the finding. Forty-plus configurations across four model families, and not one was better on quality, latency, and cost at once. Be intentional with your tradeoffs.
- A specialized model moves the frontier. On a narrow task, trained on your own data, it can beat every off-the-shelf option on all three axes even across frontier releases.
Ship on evidence, not vibes.
Replay your own traffic through every model you are considering, or talk it through with the team.