LanguaTalk

Every model was a tradeoff. So LanguaTalk built their own.

Insights fromDon PottingerCo-founder & CTO, LanguaTalk
The custom SLM vs Sonnet 5
7x faster

282ms vs 1,943ms at p50

Same quality

graded by LanguaTalk's own calibrated judge

3x lower cost

per request

The assumption

LanguaTalk's flagship product, Langua, gives 30,000+ learners a 24/7 AI conversation partner in 20+ languages. Corrections are one of its most-used AI features: the learner writes a sentence in Spanish, and the app writes it back with the mistake struck through and the fix in bold. Many learners run it automatically, so a fresh correction lands after nearly every message. It has to land fast. A correction that shows up after the conversation has moved on feels intrusive rather than helpful.

Langua's chat UI showing a live Spanish correction, with the wrong word struck through and the fix in bold
Built for LanguaTalk's Spanish corrections feature. The learner writes a sentence, and the app writes it back with the mistake struck through and the fix in bold.

The feature ran on Claude Sonnet 4, and like most teams shipping AI features, LanguaTalk assumed the next model would be the better model. Then they tried Sonnet 4.5. It behaved differently with their prompts and was worse in the product, so they rolled back to Sonnet 4. When Sonnet 4.6 came out, the team tested it by hand, working through known failure behaviours one at a time, and moved to it.

Manual testing caught real problems, but it did not scale. The team was starting to run evals, and that is when three.dev came in. What they needed next was process and tooling: a way for the team to agree on what "good" looks like, a way to test new models and prompts quickly and reliably, and a full picture of how each candidate performed on quality, latency, and cost.

Using three.dev has been a huge leap for us in providing our users the highest quality while maintaining fast responses.

Don Pottinger, Co-founder & CTO, LanguaTalk

Agree on what "good" means

Corrections have no click-through rate. Nothing in the product says a correction was right, and users are not a reliable source of high-volume feedback on quality.

So LanguaTalk turned to an AI Judge: a model whose only job is to grade the feature's outputs, answering one question on every request: "Is this output acceptable to show to a user?"

They had tried LLM-as-a-judge before and were skeptical, with good reason. Without a way to calibrate the judge and measure its alignment, you are tricking yourself into a false sense of confidence.

We tried using LLM judges before, I just don't have a lot of confidence in what it's returning. Without some way to measure and improve alignment, previous evals were hard to trust.

Don Pottinger, Co-founder & CTO, LanguaTalk

That skepticism is exactly what three.dev's calibration process exists to answer. The judge was measured against pass/fail assessments from LanguaTalk's own domain experts and refined until it agreed with them more than 80% of the time. A judge they could trust is what made the experiment results trustworthy.

The judge earns trust by agreeing with LanguaTalk's experts more than 80% of the time.

Experts
Mark pass or fail
↓
Rubric
Refined where they disagree
↓
Judge
Re-scored on the same verdicts

The loop repeats from step three back to step one.

80%+

Alignment
Calibration is built into the platform: your domain experts grade a sample, three.dev refines the judge where it disagrees with them, and the alignment score tells you how far to trust every result that follows.

Try everything

With a judge they could trust, LanguaTalk no longer had to ask whether the newest Sonnet beat the last one. They could ask the real question: what is the best possible setup for this feature?

The bar was never quality alone. A correction has to land while the learner is still reading, so every candidate has to meet the same requirement the product lives under: a median (p50) latency of about one second, at a cost that makes sense for a feature that fires on nearly every message.

three.dev replayed their real historical traffic through 40+ variants spanning Anthropic, OpenAI, Google, and open-weight models: different models, prompts, reasoning levels, thinking on and off, and temperature. Each variant answered the same conversations, so the comparison was apples to apples.

Every single result was a tradeoff.

40+ variants, 4 model families, 0 free upgrades.

Better →← Cheaper7% worsein production+19%+44%
+44%
gpt-5.6-luna low
+19%
gemini-3.5-flash low
7% worse
gpt-5.6-luna none
In production
sonnet-4
shaded · dominated
Labeled points are measured: quality is the delta against LanguaTalk's Sonnet 4 setup on their own calibrated judge. Cost positions and the unlabeled points are schematic, drawn to show the shape of the result.
GPT-5.6 Lunareasoning low
Quality, vs Sonnet 4 (their setup at the time)+44%best of any variant
The catch2.4s p50over budget
Gemini 3.5 Flashreasoning low
Quality, vs Sonnet 4 (their setup at the time)+19%
The catch3.8s p50over budget
GPT-5.6 Lunareasoning none
Quality, vs Sonnet 4 (their setup at the time)7% worse
The catchCheap and fast, and not high enough quality

Inside the latency and cost the product needed, nothing beat their setup on quality. Every model that beat it on quality was too slow or too expensive to ship.

Build the model you want

If no existing model sits where you need it, the remaining move is to make one. LanguaTalk used three.dev to create a small language model (SLM), fine-tuned on their own corrections data and specialized for exactly one job: Spanish corrections, in their format, aligned with their judgment.

The SLM did not land on the frontier. It moved it. Against Sonnet 4.6, the model in production by then, the SLM won on all three axes: better quality (+13%, with the fewest regressions of any variant tested), 282ms at p50, and 56% lower cost per request.

282ms: the correction lands while the learner is still reading.

Better →← Faster282ms1,009ms1,487ms1,943ms
282ms
the custom SLM
1,009ms
gpt-5.6-terra none
1,487ms
sonnet-4.6
1,943ms
sonnet-5 low
dashed · 1s p50 budget
Quality vs Sonnet 4.6. GPT-5.6 Terra sits right at the one-second line. The SLM matches its quality at 3.6x the speed.

What happens when the next frontier model ships

The obvious worry about a specialized model: it is tuned to today, and the labs ship every few months. So that got measured too. Each notable frontier release goes through the same replay, against the same judge, head-to-head with the SLM.

Measured on 1,464 real conversations, every variant answering the same inputs, graded by LanguaTalk's calibrated judge. Quality and cost are relative to Sonnet 4.6, the model the SLM replaced.

Sonnet 4.6the model the SLM replaced
Qualitybaseline
p50 latency1,487ms
Cost/requestbaseline
Sonnet 5reasoning low, thinking off
Quality+11%
p50 latency1,943ms
Cost/request+46%
GPT-5.6 Terrareasoning none
Quality+14%
p50 latency1,009msright at the budget line
Cost/request−20%
The custom SLM
Quality+13%
p50 latency282ms
Cost/request−56%

On quality, the SLM, Sonnet 5, and GPT-5.6 Terra are a statistical tie. On speed it is not close: the SLM is 5.3x faster than the Sonnet 4.6 it replaced, 6.9x faster than Sonnet 5, and 3.6x faster than GPT-5.6 Terra, comfortably inside the one-second budget. GPT-5.6 Terra sits right at that line, and every Sonnet configuration is over it. On cost, the SLM is the cheapest option on the board, at roughly a third of Sonnet 5's bill per request.

The reason is structural. A frontier model has to be good at everything. A model trained on your own data only has to be good at one thing.

The next frontier release shipped. The SLM is still the best config on the board.

vs sonnet-4.6 · replacedvs sonnet-5 · next release

Quality

+13% better

→

Same quality

p50 latency

5.3x faster

→

6.9x faster

Cost per request

56% cheaper

→

70% cheaper

Sonnet 5 closed the quality gap and paid for it in latency and cost, so the SLM's edge moved from quality to speed and price. Same 1,464 conversations, same calibrated judge.

Sonnet 5 costing 46% more than Sonnet 4.6 was a surprise, because the two share the same list price. What changed is the tokenizer, the dictionary a model uses to chop text into the chunks you are billed for. Sonnet 5 turns the same request into about 1.4x as many chunks, so the identical request costs about 40% more even at identical per-token pricing.

This was an important learning for understanding the real cost per task. Two models can have the exact same cost per token, but what a request actually costs is task and model dependent. The only way to truly know is to run your own requests through the model and measure.

Same score, different failure modes

Even when quality ties, the failure modes underneath differ. three.dev groups every failed correction into named failure modes, so LanguaTalk could see not only how often each model failed, but how.

Sonnet 5 and GPT-5.6 Terra fail the same way. They tend to "fix" Spanish that was already valid, and they break the correction markup more often. The SLM fails the opposite way. It sometimes praises a sentence that still has an error, and it almost never breaks the markup. Same score, different personalities.

That turned the choice into a product decision. Correcting text that was already right erodes a learner's trust faster than missing an error, so LanguaTalk set over-correction as the more severe failure mode and accepted the SLM's under-correction as the failure to live with.

Same score, different failure modes.

The learner writes

Ayer yo fui al mercado y compro unas manzanas.

compro is the real error: it should be compré. The dotted yo fui is correct Spanish and should be left alone.

Over-corrects
sonnet-5 low

Rewrites Spanish that was already valid.

Over-corrects
gpt-5.6-terra none

Same habits as Sonnet 5, slightly less often.

Under-corrects
the custom SLM

Praises a sentence that still has an error.

All three land on the same quality score. What differs is the mistake underneath it. LanguaTalk chose under-correction as the failure to live with: correcting text that was already right erodes a learner’s trust faster than missing an error.

Similar overall quality scores, but each model has its own failure modes.

Where this leaves them

LanguaTalk went from manual testing to a repeatable check: any change to any part of their AI setup, whether model, prompt, or reasoning level, is replayed against real traffic and scored on quality, latency, and cost before a learner ever sees it. Integration was a handful of lines of code.

And they own the model running one of their most-used features, at a combination of quality, latency, and price no provider offers off the shelf. No lab's release schedule, deprecation calendar, or price change decides what their product does next.

What to take from it

  1. "Newer model, better product" is a guess. Overall quality can go up while the failures your users feel most get worse. Knowing which failures matter takes your team's expertise.
  2. The tradeoff is the finding. Forty-plus configurations across four model families, and not one was better on quality, latency, and cost at once. Be intentional with your tradeoffs.
  3. A specialized model moves the frontier. On a narrow task, trained on your own data, it can beat every off-the-shelf option on all three axes even across frontier releases.

Ship on evidence, not vibes.

Replay your own traffic through every model you are considering, or talk it through with the team.

Ship your next change on evidence.

Your next agent change deserves better than “LGTM, ship it.”