oathe / benchmark

The Agent Reliability Benchmark

Reliability under injected failure: what happens to an obligation after process death, compaction, a duplicate wake, stale authority, or a false "done".

Coming soon.
methodology v0 evidence published per cell results publish with the repo reds publish too
$ oathe test my-agent
ships with the repo · GitHub →

What we measure

AGENT BENCHMARK
input
model
answer
OATHE BENCHMARK
company obligation
agent allocation
failure injection
recovery
verified outcome

What a cell reports

recovery rate
work claims settled correctly after injected process death
orphans after terminal work
attempts still holding state after their claim is terminal
duplicate effects per 1k injected failures
external actions fired more than once across kill/replay cycles
false completions caught
"done" claims failing verification against the pre-run contract