Labs
Using semantic benchmarks to build a self-improving text-to-query agent
Our text-to-query agent went from 45 seconds on frontier models to 2 seconds on GLM 5.3 Flash, with the same accuracy at 1/20 the cost. This post describes the architecture, evaluation loop, and lessons learned.