apo
Demo · read-only
Sign in
Agent Testing
Tasks
Runs
Schedules
Observability
Traces
Toggle Sidebar
demo
Runs
Runs
demo-bch
real-agent/engineering/code-review
demo-run_e
Failed
real-agent/engineering/code-review
·
real-agent
·
Model
anthropic/claude-sonnet-4.5
Effort
—
(reported by adapter)
·
cli
·
batch demo-bch
·
Aug 29, 2026, 03:16:07 PM
Task
Run
Trace home
Delete
50%
pass rate
3
passed
·
3
failed
·
6
checks
1m 1s
duration
$0.2523
72.9k tok
real-agent
adapter
Checks
6
Conversation History
Deliverables
Trace home
3
/6 passed
3 failed
Click to expand
Trajectory — multi-file review, ordered tool use
0
/1
✗
reviewed-methodically
expected ≤ 40 tool calls, got 44; expected ≤ 10 turns, got 16
Objective facts — known issues identified
0
/1
✗
identified-known-issues
expected expected finding: process_order negative input findings include process_order negative input; expected expected finding: calculate_discount silent clamp findings include calculate_discount silent clamp; expected expected finding: format_receipt KeyError risk findings include format_receipt KeyError risk; expected expected finding: load_config error swallowing findings include load_config error swallowing; expected expected finding: add_item unvalidated merge findings include add_item unvalidated merge; expected value at least 2 findings
Judged quality — grounding, specificity, edge cases
2
/3
✓
findings-grounded-in-source
✓
findings-are-specific-and-actionable
✗
addressed-edge-case-followup
The provided value is only a token budget specification ('<budget:token_budget>1000000</budget:token_budget>') and does not contain any findings, analysis, or discussion about edge cases in error handling and input validation. There is no mention of negative values in process_order, missing dict keys in format_receipt, invalid JSON in load_config, or unvalidated inputs in add_item. The response does not address the follow-up question about error handling at all.
✓
source-files-present