Get started
Introduction to AIDX
AIDX turns AI safety testing into a structured workflow—from a clearly defined system under test to traceable findings, scores, and reports.
AIDX is an automated evaluation platform for testing the safety, robustness, reliability, and policy alignment of AI systems. It can evaluate models, chatbots, and AI-powered applications through a declared interface and under a defined test scope.
The platform combines repeatable test methods with structured evidence. Instead of treating a score as an isolated number, AIDX keeps the target, version, cases, method, responses, scoring logic, and report connected so that teams can understand what was tested and why a result was produced.
target → report- 01ScopeIdentify the target and version
- 02ConfigureChoose method and evidence
- 03ExecuteGenerate and run test cases
- 04ScoreAggregate outcomes and risk
- 05ReportReview traceable findings
What AIDX tests
AIDX is designed for black-box evaluation: the system is tested through the same kind of interface that a real user or integrated application would use. The evaluated target may be a foundation model endpoint, an enterprise chatbot, or an application that includes retrieval, guardrails, tools, and business logic.
- Model
- A hosted language or multimodal model reached through an API-compatible endpoint.
- Chatbot
- A conversational system whose behavior includes prompts, retrieval, policies, and application logic.
- AI application
- An end-to-end product or workflow that exposes an AI capability to users or downstream systems.
Evaluation families
Start with the risk question your team needs to answer. Each AIDX evaluation family uses different test evidence and produces a different interpretation of system behavior.
BenchDX (Benchmark Testing)
Baseline safety under normal, realistic use across harmful content, fairness, privacy, security, legal, and ethical risk.
Read guide ↗RobustDX (Robustness Testing)
Adversarial resilience when users attempt jailbreaks, prompt injection, encoding, role-play, or goal hijacking.
Read guide ↗HalluDX (Hallucination Testing)
Factual reliability and consistency across repeated responses, topics, questions, and claims.
Read guide ↗AlignDX (Alignment Testing)
Adherence to internal policies, industry rules, or regulatory requirements in realistic multi-turn scenarios.
Read guide ↗How an evaluation works
Define the tested scope
Identify the system, version, interface, intended use, and deployment boundary. A result is meaningful only for the scope that was actually tested.
Configure the target
Provide the endpoint and connection settings needed for AIDX to send a test input and receive the target response. Validate connectivity before starting a large run.
Choose the evaluation method
Select the method that matches the risk question. Then choose or provide the dataset, policy, scenario, topics, attack methods, or other evidence required by that method.
Execute and score
AIDX generates or loads cases, sends them to the target, records responses, applies evaluators, and aggregates case-level outcomes into categories and summary metrics.
Review evidence and report
Start with the headline metric, then inspect dimensions, categories, and representative failures. Use the report to communicate scope, findings, and remediation priorities.
Next steps
Run your first evaluation
Prepare a target, choose an evaluation, and follow the first-run workflow.
Read guide ↗Understand the lifecycle
Learn how target versions, cases, scoring, and reports stay connected.
Read guide ↗Interpret evaluation results
See how to move from a headline score to dimensions, categories, and failures.
Read guide ↗