Labs
Building a data flywheel for our agent harness
An observability pipeline that reads production traces at corpus scale, names recurring failure modes, and proves the fix held.
Comparisons
Research, experiments and benchmarks from the team building the platform.
Labs
Our text-to-query agent went from 45 seconds on frontier models to 2 seconds on GLM 5.3 Flash, with the same accuracy at 1/20 the cost. This post describes the architecture, evaluation loop, and lessons learned.
Labs (3)
Labs
An observability pipeline that reads production traces at corpus scale, names recurring failure modes, and proves the fix held.
Labs
A median request carried 150,000 tokens. Most of our context work since comes down to one move, applied repeatedly.
Labs
How and why we made our harness model-agnostic, and what per-turn portability actually cost us.