Evaluations
Evaluation on AgentGraph is anchored to the success criteria you define when creating your Agent Request. Clear, measurable criteria are what turn "the agent seems good" into an objective delivery bar.
Success Criteria
- Defined by the Buyer during Agent Design and refined with the Builder at kickoff.
- Describe what the agent must do, with what inputs, and to what quality bar — not how it is implemented.
- Serve as the acceptance checklist at Delivery: the Buyer confirms completion once the agent meets them.
Test Cases
- Good success criteria break down into concrete test cases: a scenario, the input, and the expected behavior.
- Review test results with your Builder during development rather than only at the end — iterative evaluation catches gaps early.
Human in the Loop
- Agent quality is not one-and-done. Plan for a feedback loop where the Buyer exercises the agent on real work and the Builder tunes prompts, tools, and integrations in response.