LAB / ACTIVE EXPLORATION
Research.
We study where AI systems break, then build small tests and tools that make those failures legible.
CURRENT WORK / 01
Agent evaluation
for data work.
Our first lab effort is Data Lab: isolated Python data-pipeline tasks, hidden executable checks, oracle/no-op calibration, and repeated trials. The starter suite is synthetic; the next step is private tasks grounded in partner failures.
Explore Data LabRESEARCH POLICY / 02
Evidence first.
We publish a result only with its task definition, evaluation conditions, and limits. No papers or general model rankings are claimed here yet. If you have a failure worth studying together, send the reproduction.
Discuss a collaboration ↗