Skip to content

Getting started

Using validia

uv add validia
import validia

print(validia.__version__)

The validia command

validia init                      # validia.toml and an example suite
validia create                    # a suite of your own: prompt, grader, cases
validia add evals/ticket-triage/suite.toml       # add a test case, one question at a time
validia check evals/ticket-triage/suite.toml outage-login --reply "Urgent."
validia config                    # every setting, and where its value came from
validia run evals/ticket-triage/suite.toml -m claude-sonnet-5-5 --dry-run
validia run evals/ticket-triage/suite.toml -m claude-sonnet-5-5 -n 3   # call the model

init asks three questions -- the folder for suites, the answer type, and the example suite's name -- and Enter keeps each default. --evals DIR, --answer TYPE and --name NAME answer them up front, and without a terminal (CI, scripts) the defaults are used without asking. An existing validia.toml is kept, so running init again adds another suite:

validia.toml
evals/
└── ticket-triage/
    ├── suite.toml      # the grader and the cases, with their expected answers
    ├── prompt.md       # the prompt under test; the model never sees suite.toml
    └── validia.toml    # optional: settings for this suite alone

Creating a suite

validia create builds a suite of your own from questions, with numbered options wherever there is a choice:

$ validia create
Folder for eval suites [evals]:
Suite name: agent-call
What does the prompt answer with?
  1  label  a category from a fixed list
  2  json   a JSON object
  3  text   free text
  4  tool   a call to one of its tools
Choose 1-4 [1]: 4

It then asks for what grading needs -- the labels, the keys every JSON reply must have, or the tools (described on the spot, or an existing JSON file) -- and for the prompt: built from questions, typed in, an existing file the suite points at (so it tests the prompt you ship), or a placeholder to fill in later.

Building the prompt asks one question per section -- who the model is, its job, what it needs to know, its rules with their reasons, and for free text, its tone -- then offers a closing "how to answer" instruction that fits the answer type and stays within what the grader accepts. As you type, core checks point out weak instructions without blocking them:

A rule to follow, with its reason (empty to finish): NEVER guess
  hint caps: capitals add pressure, not meaning: say it plainly, and say why
  hint reason: give the reason (because ..., so that ...): a rule with its reason covers cases it never names
  hint forbid-only: say what to do instead: a rule that only forbids leaves the model guessing what is wanted
Keep it as it is? [y/N]:

The sections, the checks and the closing instructions are all data in a TOML template. validia template > prompt.toml saves the built-in one into the project, and from then on create builds from that file: change a question, add a check, reword an instruction. Then the cases, one at a time, in the answer type's shape. Nothing is written until the last answer, and only if the suite loads; --evals, --answer and --name answer their questions up front.

From code, or a REST API

Everything the command does is a function in the validia package, with no terminal attached: the console is one front end, and a REST service, a web form or a notebook can be another. Specs come from plain JSON, results go back with dataclasses.asdict or describe_suite, and every failure is a SpecError whose problems list is ready to return as a 422.

from dataclasses import asdict
from pathlib import Path

from fastapi import FastAPI, HTTPException  # any web framework; not a validia dependency

import validia

app = FastAPI()
PROJECT = Path(".")


@app.post("/suites")
def create(body: dict) -> dict:
    try:
        suite = validia.create_suite(validia.SuiteSpec.from_dict(body), PROJECT)
    except validia.SpecError as exc:
        raise HTTPException(422, exc.problems) from exc
    return validia.describe_suite(suite, PROJECT)


@app.post("/suites/{name}/cases")
def add(name: str, body: dict) -> dict:
    try:
        case = validia.CaseSpec.from_dict(body)
        suite = validia.add_case(validia.locate_suite(PROJECT, name), case)
    except validia.SpecError as exc:
        raise HTTPException(422, exc.problems) from exc
    return validia.describe_suite(suite, PROJECT)


@app.post("/suites/{name}/cases/{case_id}/check")
def check(name: str, case_id: str, body: dict) -> dict:
    try:
        suite = validia.load_suite(validia.locate_suite(PROJECT, name))
        verdict = validia.check_reply(suite, case_id, validia.parse_reply(body))
    except validia.SpecError as exc:
        raise HTTPException(422, exc.problems) from exc
    return asdict(verdict)


@app.get("/prompt-questions")
def questions(answer: str = "label", labels: str = "") -> list[dict]:
    grade = validia.Grade(answer, labels=tuple(filter(None, labels.split(","))))
    return [asdict(q) for q in validia.prompt_questions(validia.default_template(), grade)]

Take a suite's name from a URL through locate_suite, never by joining paths: it refuses a name such as .. that would reach outside the project, and create_suite likewise refuses a prompt_file or tools_file that leads outside it, symlinks included.

A form renders the Question objects -- each has an id, text, kind (text, many or choice), options and an example -- review_answer returns the core checks' hints for one answer as it is typed, and render_prompt turns the submitted answers into the same prompt the console builds.

Answer types

A suite's [grade] type says what the prompt answers with, and each case's expected takes the matching shape. Every check is code, and an empty reply always fails.

Type The reply is A case's expected
label a category from labels "urgent"
json a JSON object { category = "bug", priority = "high" }; required in [grade] lists keys every reply needs
text free text { contains = ["Export"], not_contains = ["Import"], matches = '(?i)zip' }, and equals
tool a call to one of the suite's tools { tool = "lookup_order", args = { order_id = "48213" } }, or { tool = "none" }

A tool suite names its tools in a JSON file (tools = "tools.json") with name, description and parameters (a JSON Schema), the same fields franca's ToolDef takes. init writes one example per type.

Adding and checking cases

validia add SUITE asks for a case's id, input and tags, then for the expected answer in the suite's shape -- a label, one value per JSON field, text checks, or a tool and its arguments -- appends it to suite.toml, and lets you try replies against it. A case that would break the suite is never written.

validia check SUITE CASE grades one reply you supply, without calling a model: the quickest way to test a regex or a contains list before any run. It exits 0 when the reply passes and 1 when it fails. The reply comes from --reply, or --tool and --args for a tool call, and otherwise from the terminal or standard input.

Checking prompts: validia lint

validia lint checks prompt files, or whole suites -- their prompt and every tool description -- against the rules for a model, without calling it:

$ validia lint evals/billing/prompt.md -m claude-sonnet-5-5
evals/billing/prompt.md:2:1  info   wording/capitals  'NEVER'  State the one real constraint plainly, ...
evals/billing/prompt.md:2:1  warn   wording/rule-without-reason  'NEVER'  Give the reason (because ..., so that ...): ...
evals/billing/prompt.md:3:1  error  reasoning/show-reasoning  'Show your reasoning'  Remove it: these models refuse ...
rules 1.0.0 for anthropic:claude-sonnet-5-5: 3 findings (1 error, 1 warn, 1 info)

It exits 1 when a finding reaches --fail-on (lint.fail_on). Which rules apply, and how severe each is, depends on the model, the provider that serves it and the release of the rules you pin. Prompt rules covers all three, how to add rules of your own, and lists every rule validia ships.

Models and API keys

model is provider:model, as in anthropic:claude-sonnet-5-5. The provider may be left out when the name says it: claude-, gpt-, gemini-, grok- and deepseek-.

API keys never go in a settings file. They are read from the environment when a run calls the model: the variable api_key_env names, then FRANCA_<PROVIDER>_API_KEY, then the vendor's own (ANTHROPIC_API_KEY, OPENAI_API_KEY, ...). The rest of a provider's access is franca's own [providers.<name>] table:

[providers.anthropic]
api_key_env = "TEAM_ANTHROPIC_KEY"   # the variable that holds the key; never the key
base_url = "https://gateway.example/v1"
timeout_s = 60

The simplest place for keys is a .env file in the project, which .gitignore already keeps out of the repository:

ANTHROPIC_API_KEY=...
OPENAI_API_KEY=...
GEMINI_API_KEY=...

A key exported in the shell outranks the same key in .env, and an empty variable counts as unset. validia config --keys lists every provider and where its key comes from -- anthropic key from ANTHROPIC_API_KEY in .env:2 -- or what to set, and never prints a key itself. run --dry-run does the same for the model it would call, before anything is spent.

Per-suite settings

A validia.toml in a suite's folder outranks the project's for that suite, so one suite can use another model, more reps or a different provider. validia config --suite SUITE shows settings as that suite sees them.

Running a suite: validia run

validia run sends every case to the model, --reps times each, and grades every reply with the suite's own check -- the same one validia check uses:

$ validia run evals/ticket-triage/suite.toml -m claude-sonnet-5-5 -n 3
ticket-triage on anthropic:claude-sonnet-5-5: 18 of 24 passed, 75.0% (95% interval 55.1%-88.0%)
  served by claude-sonnet-5-5-20261001
  by group: urgent 6/12, normal 12/12
  failing:
    data-loss-invoices      0 of 3  expected 'urgent', got 'normal'
    security-unknown-login  0 of 3  expected 'urgent', got 'normal'
  tokens: 2,880 in, 48 out
  latency: p50 820 ms, p95 1,900 ms
  written to .validia/runs/20261007T143012Z-ticket-triage/

The suite's prompt is the system text and each case's input the user turn. Repetitions make the pass rate an estimate, so it comes with a 95% interval: with eight cases and one rep, 75% means "somewhere between 41% and 93%". A call that fails -- a rate limit that outlives --retries, a bad key -- is counted as an error, never as a wrong answer, and fails the run. A failure every trial would hit the same way -- a bad key, a retired model -- stops the run at the first one: calls already in flight finish, and nothing more is sent. --fail-under 90 (run.fail_under) also fails it when the pass rate is below 90%, for CI. Every trial -- reply, tokens, latency, attempts, served model -- is written to trials.jsonl, and the totals to summary.json, in a folder of the run's own under run.output.

Calling a model needs an HTTP client: pip install 'validia[http]'. The key comes from the environment or the project's .env, as validia config --keys shows. --dry-run calls nothing: it validates the suite, reporting every problem in it at once, checks model access, and prints what would run. Suites with tools are refused for now: franca, the layer validia calls models through, sends text turns only in its current release, so the tools would never reach the model.

Settings resolve through whence. Highest precedence first:

Source Example
A command's own flags validia run suite.toml --reps 3
--set KEY=VALUE validia config --set run.reps=3
Environment variables VALIDIA_RUN__REPS=3 (__ separates sections)
.env in the working directory VALIDIA_RUN__REPS=3
A settings file --config FILE or $VALIDIA_CONFIG, then the suite folder's validia.toml, then the project's
pyproject.toml a [tool.validia] table
Defaults what validia init writes

-p NAME activates a profile, overlaying validia.NAME.toml on validia.toml. A misspelled key or an out-of-range value fails with the file it came from and a suggestion, and validia config --discovery lists every place that was searched.

Local development

The repository is a single package managed with uv (0.12 or newer). Every command below runs from the repository root.

git clone https://github.com/izmailov-labs/validia
cd validia
make install     # sync dev + docs groups, install pre-commit hooks
validia/
├── pyproject.toml        # package metadata and the lint, type, test and coverage gate
├── src/validia/          # the package on PyPI
│   ├── suites/           # suites, graders, suite files, example suites, the suite API
│   ├── prompts/          # the guided prompt builder and its template
│   ├── rules/            # the lint engine, the core rules (core/), project rule files
│   ├── runs/             # model access, and running a suite against a model
│   └── cli/              # the validia command, its questions and its settings
├── tests/
└── docs/                 # this site

validia depends on two sibling packages, franca and whence, which resolve from PyPI like any other dependency.

Common tasks:

Command What it does
make lint ruff check + ruff format --check
make fmt Autofix and format
make typecheck mypy in strict mode
make test Run the test suite
make cov Tests with coverage (fails under 90%)
make docs Build these docs with --strict
make build Build the sdist + wheel and validate metadata
make all Everything CI runs

Tests run under pytest-asyncio in auto mode, so async def test_* functions need no decorator.