Agentic evaluation
Agentic evaluation
Section titled “Agentic evaluation”This page is generated from the live QualityMax MCP registry. Examples use placeholders; replace them with values for your workspace.
export_agentic_eyes_report
Section titled “export_agentic_eyes_report”Purpose
Section titled “Purpose”Generate a standalone HTML or PDF report from a saved, project-attributed Agentic Eyes persona review. HTML is returned as text; PDF is returned as base64 so MCP clients can save it without binary corruption.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
project_id |
integer |
Yes | QualityMax project that owns the report. |
run_id |
string |
Yes | Persisted Agentic Eyes run id returned by run_persona_review. |
format |
string |
No | Export format. PDF content is base64 encoded. |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": { "project_id": 1, "run_id": "your-run-id" }, "tool": "export_agentic_eyes_report"}get_agent_eval_results
Section titled “get_agent_eval_results”Purpose
Section titled “Purpose”Get detailed results for an agent evaluation run, including per-test scores, judge reasoning, and conversation transcripts.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
run_id |
string |
Yes | Evaluation run ID |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": { "run_id": "your-run-id" }, "tool": "get_agent_eval_results"}list_adversarial_prompts
Section titled “list_adversarial_prompts”Purpose
Section titled “Purpose”Browse the adversarial prompt library. 39 curated attack prompts across 6 categories for testing AI agent safety.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
category |
string |
No | Filter by category (optional) |
severity_min |
string |
No | Minimum severity: low, medium, high, critical |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": {}, "tool": "list_adversarial_prompts"}list_end_user_personas
Section titled “list_end_user_personas”Purpose
Section titled “Purpose”List every built-in Agentic Eyes end-user persona plus the authenticated user’s saved custom personas. Each result includes a runnable_id that can be passed directly as the persona argument to run_persona_review or run_persona_consensus.
Parameters
Section titled “Parameters”This tool takes no arguments.
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": {}, "tool": "list_end_user_personas"}run_adversarial_eval
Section titled “run_adversarial_eval”Purpose
Section titled “Purpose”Run adversarial safety testing against an AI agent. Sends 39 attack prompts (prompt injection, jailbreak, data extraction, PII leakage, bias probing, off-topic) and reports which attacks succeeded vs were blocked.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
project_id |
integer |
Yes | Project ID |
agent_endpoint |
string |
Yes | Agent API endpoint URL |
categories |
array |
No | Filter by category: prompt_injection, jailbreak, data_extraction, pii_leakage, bias_probing, off_topic |
severity_min |
string |
No | Minimum severity: low, medium, high, critical (default: low) |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": { "agent_endpoint": "your-agent-endpoint", "project_id": 1 }, "tool": "run_adversarial_eval"}run_agent_eval
Section titled “run_agent_eval”Purpose
Section titled “Purpose”Run a conversation evaluation against an AI agent endpoint. Sends test case prompts to the agent, judges responses on accuracy/helpfulness/safety/relevance/conciseness. Returns per-test scores and overall pass rate.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
project_id |
integer |
Yes | Project ID containing agent_eval test cases |
agent_endpoint |
string |
Yes | Agent API endpoint URL (e.g. https://my-app.com/api/chat) |
criteria |
array |
No | Evaluation criteria (default: accuracy, helpfulness, safety, relevance, conciseness) |
pass_threshold |
number |
No | Minimum score to pass (0-1, default: 0.7) |
test_case_ids |
array |
No | Specific test case IDs to evaluate (omit for all agent_eval cases) |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": { "agent_endpoint": "your-agent-endpoint", "project_id": 1 }, "tool": "run_agent_eval"}run_persona_consensus
Section titled “run_persona_consensus”Purpose
Section titled “Purpose”Run the SAME persona review across several different AI models and rank the findings by how many models independently reported them. Findings that agree across models are high-confidence (‘4/4 models flagged no pricing’); findings only one model saw are exploratory and listed in a divergence section. Use this when a finding is going in front of a customer or driving a fix, and a single model’s word isn’t enough. Costs roughly N times a single review. Part of Agentic Eyes.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
url |
string |
Yes | Live URL to review (http/https). |
persona |
string |
No | Which end-user persona browses the site: a builtin slug (sally, worst_customer) or the UUID of a persona saved via the persona library. |
models |
array |
No | Model ids to compare (2-5). Omit to auto-pick a cross-provider panel from the verified model catalog. |
goal |
string |
No | Optional task the persona tries to accomplish (e.g. ‘book a lagoon tour’). |
drive |
boolean |
No | If true, each model drives a real browser through the flow instead of reviewing only the landing page. |
max_steps |
integer |
No | Drive mode only: cap on navigation steps (defaults to a patience-scaled budget). |
evaluation_mode |
string |
No | Explicit evaluation semantics for every panel run. Legacy omission defaults to conversion and is reported as defaulted in the consensus and child-run responses. |
project_id |
integer |
No | Optional QualityMax project to attribute the runs to. |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": { "url": "https://example.com" }, "tool": "run_persona_consensus"}run_persona_review
Section titled “run_persona_review”Purpose
Section titled “Purpose”Browse a live website as a customizable AI end-user persona (e.g. Sally, the Dreaming Planner, or the Worst-Customer-Ever) and return a plain-language ‘message from user’ UX report: what they tried, where they got stuck, what confused them, the questions they had, whether they’d convert, and the top fixes. With drive=true the persona actually clicks through a multi-step flow (not just the landing page) and reports the journey it took. Part of Agentic Eyes.
Parameters
Section titled “Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
url |
string |
Yes | Live URL to review (http/https). |
persona |
string |
No | Which end-user persona browses the site: a builtin slug (sally, worst_customer) or the UUID of a persona saved via the persona library. |
goal |
string |
No | Optional task the persona tries to accomplish (e.g. ‘book a lagoon tour’). |
drive |
boolean |
No | If true, drive a real browser through the flow (click/type) instead of reviewing only the landing page. |
max_steps |
integer |
No | Drive mode only: cap on navigation steps (defaults to a patience-scaled budget). |
allow_external_actions |
boolean |
No | Explicit authorization to click state-changing controls. Credentials and identity fields remain blocked. |
evaluation_mode |
string |
No | Explicit evaluation semantics. Use pre_signup_intent for no-credential purchase/trial research. Legacy omission defaults to conversion and is reported as defaulted in the response. |
project_id |
integer |
No | Optional QualityMax project to attribute the run to. |
Return shape
Section titled “Return shape”The live registry does not declare a response schema. Expect a tool-specific JSON result and handle success, error, and result fields defensively.
Worked example
Section titled “Worked example”{ "arguments": { "url": "https://example.com" }, "tool": "run_persona_review"}