Skip to content

CriterIAQA · New service

Would you put your customers in a car without an inspection?

Your AI agents talk to real customers, quote prices, make promises, and make decisions. CriterIAQA is roadworthiness testing for your AI agents: we inspect them before they reach production and give you a clear verdict, ready or not ready, backed by evidence, not promises.

CriterIAQA Report

AGT-0192 · Booking Assistant

Approved
  • Response testing
  • Escalation rules
  • Traceability (evidence)
  • Authorized data

Ready for production

CriterIAQA verdict · illustrative example

What it is

The quality control almost no one runs on their agents today

A car doesn't go on the road without its annual inspection. An AI agent, today, does reach production without anyone inspecting it. CriterIAQA closes that gap: we test your agent before it talks to a real customer and hand you a verdict, not an opinion.

Without QA

The agent goes straight from development to production. No one knows how it responds under pressure, whether it hallucinates, whether it follows business rules, or whether a customer can manipulate it. The first real test is an upset customer or a complaint.

With CriterIAQA

The agent goes through a 4-point checklist, with documented evidence for each test. You come out with a clear verdict: green, yellow, or red. If needed, you also get a concrete correction plan before launch.

How it works

The CriterIAQA 4-point checklist

The checklist is grounded in recognized practices for evaluating agents: task completion, tool use, traceability, and instruction-following. We apply it with human judgment and evidence for your business.

Response testing

Does the agent finish the task, or stall halfway through? We test the full flow, not just the happy path.

Escalation rules

Does it use the right tools, or fail invisibly? We validate when it should resolve on its own and when it should escalate to a human.

Traceability (evidence)

Can every response be traced back to a real source, without hallucinating? We document the origin of each claim the agent makes.

Authorized data

Does it respect the rules and data it's actually allowed to use? We verify it doesn't expose or invent what it shouldn't.

What you get

A clear verdict: ready or not ready for production

At the end of the inspection you get a readiness signal, not a 40-page report you have to interpret on your own.

Green

Ready for production. The agent passed the 4-point checklist with documented evidence.

Yellow

Fix before launch. We identify the exact points to adjust and the plan to reach green.

Red

Not fit for real customers. The agent should not go to production in its current state.

CriterIAQA Report: every agent evaluated gets documented evidence, ready to show leadership, legal, or your own end client.

Why it matters

When no one ran QA before production

These cases have been publicly reported. None involved cutting-edge technology: a basic checklist before going to production is directly related to what went wrong.

Air Canada

Feb. 2024

The chatbot invented a refund policy that didn't exist. A tribunal ordered the airline to honor it.

Related control: Traceability

Chevrolet (dealership)

Dec. 2023

A chatbot agreed to sell a car for US$1 after a prompt-injection attack.

Related control: Response testing

DPD

Jan. 2024

A user manipulated the chatbot into insulting its own company. It went viral.

Related control: Response testing

Lawyer vs. Avianca (U.S.)

Jun. 2023

A lawyer submitted court filings citing legal cases invented by a generative AI. They never existed.

Related control: Traceability

The risk of skipping it

Why do QA, and what happens if you don't

The same cases, seen through the risk that could have been avoided.

Binding legal liability

A tribunal ordered Air Canada to honor a policy its own chatbot invented.

CriterIAQA aims to catch failures before they reach the customer, not after the crisis, and gives you a traceable, defensible basis before leadership, legal, and regulators.

Viral reputational damage

DPD's chatbot insulted its own company on social media. It went viral.

Turns trust into a verifiable process, not a personal bet by whoever approves the launch.

Direct financial loss

A dealership chatbot "sold" a car for US$1 after a prompt attack.

A basis for moving from pilot to production with clear exit criteria, not the hope that no one tests it.

Legal exposure from hallucinations

A lawyer submitted court filings citing legal cases invented by a generative AI.

We document evidence of control over the agent's responses before they reach a customer or a court.

Frameworks like the NIST AI RMF, ISO/IEC 42001 and, where applicable, the EU AI Act help structure controls, traceability, and oversight. This information is general and does not replace legal advice.

Schedule your inspection

Before your agent talks to its next customer

An initial session to review your agent against the CriterIAQA checklist and define the scope of the full assessment.