“I did it.”
The system describes an action as complete even though the result cannot be found or checked.
- Claim separated from evidence
- Verification pushed onto the user
Artificial intelligence claims the future. Agent Exposed records what actually happened. Company by company. Report by report. Claim versus result.
Tell us what the agent claimed, what it actually did, and what you could verify. Screenshots, logs, diffs, and reproducible examples support review but are not required to submit a report.
The system describes an action as complete even though the result cannot be found or checked.
The system implies it inspected a file, app, source, or result without showing a reliable trace.
New code, plans, downloads, and explanations while the original objective remains unsolved.
When the agent fails, the cost gets pushed back to the user: reinstall, restart, repeat, begin again.
Published evaluations must name their source, version, date, model or agent configuration, and test conditions. A user report is not automatically a benchmark result.
| System | Benchmark | Status |
|---|
Vote on agents, providers, and models. Add a model anonymously. Let the public record show what users actually encounter.
Evidence is optional for submission. Verification is visible, earned, and never granted merely because an image was uploaded.