Skip to content

LAB / ACTIVE EXPLORATION

Research.

We study where AI systems break, then build small tests and tools that make those failures legible.

CURRENT WORK / 01

Agent evaluation
for data work.

Our first lab effort is Data Lab: isolated Python data-pipeline tasks, hidden executable checks, oracle/no-op calibration, and repeated trials. The starter suite is synthetic; the next step is private tasks grounded in partner failures.

Explore Data Lab

RESEARCH POLICY / 02

Evidence first.

We publish a result only with its task definition, evaluation conditions, and limits. No papers or general model rankings are claimed here yet. If you have a failure worth studying together, send the reproduction.

Discuss a collaboration ↗