Blog
Notes on how AI agents fail, what actually improves quality, latency, and cost, and the teams shipping on evidence instead of vibes.

Product
Your use case is not on the leaderboard
We replayed 2,000 PII-redaction requests through seven frontier models. The consensus best model improved quality but at four and a half times the cost, upgrading to a bigger model added latency and cost without improving quality, and the configuration that won on all three axes is one no public benchmark would have pointed to.
Andrea Moscatelli · 
Company
Hello, world. We're round3.
We started a company to answer the hardest question in production AI: is this change actually better?
Borja Burgos ·