apo
Demo · read-only
Sign in
Agent Testing
Tasks
Runs
Schedules
Observability
Traces
Toggle Sidebar
demo
Runs
Runs
demo-bch
real-agent/documents/document-qa
demo-run_b
Failed
real-agent/documents/document-qa
·
real-agent
·
Model
anthropic/claude-haiku-4-5
Effort
—
(reported by adapter)
·
cli
·
batch demo-bch
·
Aug 29, 2026, 03:07:48 PM
Task
Run
Trace home
Delete
40%
pass rate
2
passed
·
3
failed
·
5
checks
16.9s
duration
$0.0263
18.6k tok
real-agent
adapter
Checks
5
Conversation History
Deliverables
Trace home
2
/5 passed
3 failed
Click to expand
✓
read-spec-and-searched
✗
answers-all-five-questions
expected auth flow includes client[- ]credentials
✗
answers-grounded-in-spec
To evaluate 'answers-grounded-in-spec', I need to assess whether the agent's answers demonstrate that they actually read and extracted information from the specific technical specification document, rather than providing generic API knowledge.
✗
answers-are-complete-and-useful
To evaluate this task, I need to assess whether the agent provided five complete answers with sufficient context. However, I notice that no actual deliverable/result has been provided in the prompt for me to grade. The rubric lists five values to evaluate (API version 2.4, authentication method with JWT/OAuth 2.0, 11 endpoints across three categories, PostgreSQL 16 and Redis 7.2), but without seeing the agent's actual answers, I cannot determine if they included contextual information or were merely bare values. The instruction asks me to check if each answer is 'accompanied by enough context to be useful' — this requires examining the actual response content, which is absent from this evaluation request.
✓
spec-file-present