Summary
WhatWorkedBench is introduced to evaluate AI agents' ability to predict outcomes of experimental changes by inspecting code and forecasting performance across configurations. It uses exhaustive CPU execution to provide accurate response surfaces for component settings.
AI-assisted summary based on the listed source.
What happened
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a...
Signal Intelligence
Signal Strength 95%
Technical label SOURCE-BACKED
Public Interest 30
Category RESEARCH
Reader Depth TECHNICAL
Signal Strength reflects source quality, relevance, freshness and evidence. Public Interest helps organize discovery; it is not proof of truth.
Public Interest components
Recognizable Entity Score 0
Practical Impact Score 0
Novelty Interest Score 94
Consequence Score 34
Curiosity Score 16
Shareability Score 45