Skip to content

Agentic evaluation

This page is generated from the live QualityMax MCP registry. Examples use placeholders; replace them with values for your workspace.

Generate a standalone HTML or PDF report from a saved, project-attributed Agentic Eyes persona review. HTML is returned as text; PDF is returned as base64 so MCP clients can save it without binary corruption.

Parameter Type Required Description
project_id integer Yes QualityMax project that owns the report.
run_id string Yes Persisted Agentic Eyes run id returned by run_persona_review.
format string No Export format. PDF content is base64 encoded.

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {
"project_id": 1,
"run_id": "your-run-id"
},
"tool": "export_agentic_eyes_report"
}

Get detailed results for an agent evaluation run, including per-test scores, judge reasoning, and conversation transcripts.

Parameter Type Required Description
run_id string Yes Evaluation run ID

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {
"run_id": "your-run-id"
},
"tool": "get_agent_eval_results"
}

Browse the adversarial prompt library. 39 curated attack prompts across 6 categories for testing AI agent safety.

Parameter Type Required Description
category string No Filter by category (optional)
severity_min string No Minimum severity: low, medium, high, critical

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {},
"tool": "list_adversarial_prompts"
}

List every built-in Agentic Eyes end-user persona plus the authenticated user’s saved custom personas. Each result includes a runnable_id that can be passed directly as the persona argument to run_persona_review or run_persona_consensus.

This tool takes no arguments.

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {},
"tool": "list_end_user_personas"
}

Run adversarial safety testing against an AI agent. Sends 39 attack prompts (prompt injection, jailbreak, data extraction, PII leakage, bias probing, off-topic) and reports which attacks succeeded vs were blocked.

Parameter Type Required Description
project_id integer Yes Project ID
agent_endpoint string Yes Agent API endpoint URL
categories array No Filter by category: prompt_injection, jailbreak, data_extraction, pii_leakage, bias_probing, off_topic
severity_min string No Minimum severity: low, medium, high, critical (default: low)

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {
"agent_endpoint": "your-agent-endpoint",
"project_id": 1
},
"tool": "run_adversarial_eval"
}

Run a conversation evaluation against an AI agent endpoint. Sends test case prompts to the agent, judges responses on accuracy/helpfulness/safety/relevance/conciseness. Returns per-test scores and overall pass rate.

Parameter Type Required Description
project_id integer Yes Project ID containing agent_eval test cases
agent_endpoint string Yes Agent API endpoint URL (e.g. https://my-app.com/api/chat)
criteria array No Evaluation criteria (default: accuracy, helpfulness, safety, relevance, conciseness)
pass_threshold number No Minimum score to pass (0-1, default: 0.7)
test_case_ids array No Specific test case IDs to evaluate (omit for all agent_eval cases)

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {
"agent_endpoint": "your-agent-endpoint",
"project_id": 1
},
"tool": "run_agent_eval"
}

Run the SAME persona review across several different AI models and rank the findings by how many models independently reported them. Findings that agree across models are high-confidence (‘4/4 models flagged no pricing’); findings only one model saw are exploratory and listed in a divergence section. Use this when a finding is going in front of a customer or driving a fix, and a single model’s word isn’t enough. Costs roughly N times a single review. Part of Agentic Eyes.

Parameter Type Required Description
url string Yes Live URL to review (http/https).
persona string No Which end-user persona browses the site: a builtin slug (sally, worst_customer) or the UUID of a persona saved via the persona library.
models array No Model ids to compare (2-5). Omit to auto-pick a cross-provider panel from the verified model catalog.
goal string No Optional task the persona tries to accomplish (e.g. ‘book a lagoon tour’).
drive boolean No If true, each model drives a real browser through the flow instead of reviewing only the landing page.
max_steps integer No Drive mode only: cap on navigation steps (defaults to a patience-scaled budget).
evaluation_mode string No Explicit evaluation semantics for every panel run. Legacy omission defaults to conversion and is reported as defaulted in the consensus and child-run responses.
project_id integer No Optional QualityMax project to attribute the runs to.

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {
"url": "https://example.com"
},
"tool": "run_persona_consensus"
}

Browse a live website as a customizable AI end-user persona (e.g. Sally, the Dreaming Planner, or the Worst-Customer-Ever) and return a plain-language ‘message from user’ UX report: what they tried, where they got stuck, what confused them, the questions they had, whether they’d convert, and the top fixes. With drive=true the persona actually clicks through a multi-step flow (not just the landing page) and reports the journey it took. Part of Agentic Eyes.

Parameter Type Required Description
url string Yes Live URL to review (http/https).
persona string No Which end-user persona browses the site: a builtin slug (sally, worst_customer) or the UUID of a persona saved via the persona library.
goal string No Optional task the persona tries to accomplish (e.g. ‘book a lagoon tour’).
drive boolean No If true, drive a real browser through the flow (click/type) instead of reviewing only the landing page.
max_steps integer No Drive mode only: cap on navigation steps (defaults to a patience-scaled budget).
allow_external_actions boolean No Explicit authorization to click state-changing controls. Credentials and identity fields remain blocked.
evaluation_mode string No Explicit evaluation semantics. Use pre_signup_intent for no-credential purchase/trial research. Legacy omission defaults to conversion and is reported as defaulted in the response.
project_id integer No Optional QualityMax project to attribute the run to.

The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.

{
"arguments": {
"url": "https://example.com"
},
"tool": "run_persona_review"
}