Skip to content

API reference

validia

validia — a universal, async-first evaluation framework for LLM systems.

Covers the full evaluation spectrum under one set of primitives: prompt evaluation, tool selection, agent evaluation, and agent-type comparison.

The public API is re-exported from this module; everything not listed in __all__ is internal and may change without a major version bump. It is the same API whatever the front end: the validia command calls these functions, and so can a REST service or a notebook -- see :mod:validia.suites.api.

CATEGORIES module-attribute

CATEGORIES = (
    "wording",
    "context",
    "reasoning",
    "output",
    "tools",
    "security",
    "maintenance",
)

The core categories, in reading order.

CLOSING module-attribute

CLOSING = 'closing'

The id of the last question: how the model should answer.

PromptTemplate dataclass

A loaded, checked template.

Attributes:

Name Type Description
sections tuple[Section, ...]

The questions, in order.

hints tuple[Hint, ...]

The core checks run on each answer.

output dict[AnswerType, tuple[str, ...]]

For each answer type, the closing instructions to offer, first one the default.

source str

Where the template came from, for the user to see.

Source code in src/validia/prompts/template.py
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
@dataclass(frozen=True, slots=True)
class PromptTemplate:
    """A loaded, checked template.

    Attributes:
        sections: The questions, in order.
        hints: The core checks run on each answer.
        output: For each answer type, the closing instructions to offer, first
            one the default.
        source: Where the template came from, for the user to see.
    """

    sections: tuple[Section, ...]
    hints: tuple[Hint, ...]
    output: dict[AnswerType, tuple[str, ...]]
    source: str

    def sections_for(self, answer: AnswerType) -> tuple[Section, ...]:
        """List the sections that apply to an answer type, in order.

        Args:
            answer: The suite's answer type.

        Returns:
            The sections.
        """
        return tuple(s for s in self.sections if not s.answers or answer in s.answers)

    def hints_for(self, section: str, answer: str) -> list[Hint]:
        """List the hints an answer sets off.

        Args:
            section: The id of the section the answer is for.
            answer: The answer.

        Returns:
            The hints that fire, in template order.
        """
        return [hint for hint in self.hints if hint.fires(section, answer)]

sections_for

sections_for(answer: AnswerType) -> tuple[Section, ...]

List the sections that apply to an answer type, in order.

Parameters:

Name Type Description Default
answer AnswerType

The suite's answer type.

required

Returns:

Type Description
tuple[Section, ...]

The sections.

Source code in src/validia/prompts/template.py
157
158
159
160
161
162
163
164
165
166
def sections_for(self, answer: AnswerType) -> tuple[Section, ...]:
    """List the sections that apply to an answer type, in order.

    Args:
        answer: The suite's answer type.

    Returns:
        The sections.
    """
    return tuple(s for s in self.sections if not s.answers or answer in s.answers)

hints_for

hints_for(section: str, answer: str) -> list[Hint]

List the hints an answer sets off.

Parameters:

Name Type Description Default
section str

The id of the section the answer is for.

required
answer str

The answer.

required

Returns:

Type Description
list[Hint]

The hints that fire, in template order.

Source code in src/validia/prompts/template.py
168
169
170
171
172
173
174
175
176
177
178
def hints_for(self, section: str, answer: str) -> list[Hint]:
    """List the hints an answer sets off.

    Args:
        section: The id of the section the answer is for.
        answer: The answer.

    Returns:
        The hints that fire, in template order.
    """
    return [hint for hint in self.hints if hint.fires(section, answer)]

TemplateError

Bases: ValueError

A prompt template is malformed; the message lists every problem in it.

Source code in src/validia/prompts/template.py
56
57
class TemplateError(ValueError):
    """A prompt template is malformed; the message lists every problem in it."""

Finding dataclass

One hit of one rule.

Attributes:

Name Type Description
rule str

The rule's id.

severity Severity

How much it matters on the target model.

line int

One-based line of the hit.

column int

One-based column of the hit.

excerpt str

The text that matched, shortened.

fix str

What to do about it.

Source code in src/validia/rules/model.py
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
@dataclass(frozen=True, slots=True)
class Finding:
    """One hit of one rule.

    Attributes:
        rule: The rule's id.
        severity: How much it matters on the target model.
        line: One-based line of the hit.
        column: One-based column of the hit.
        excerpt: The text that matched, shortened.
        fix: What to do about it.
    """

    rule: str
    severity: Severity
    line: int
    column: int
    excerpt: str
    fix: str

Guidance dataclass

How to write for one target, in one category, with where the claims come from.

Attributes:

Name Type Description
category str

The category.

target str

<provider>/default or <provider>/<model>.

version str

The release its file comes from.

summary str

What matters about the target, in a sentence.

instructions tuple[str, ...]

What to do when writing a prompt for it.

sources tuple[str, ...]

Where each claim comes from.

Source code in src/validia/rules/model.py
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
@dataclass(frozen=True, slots=True)
class Guidance:
    """How to write for one target, in one category, with where the claims come from.

    Attributes:
        category: The category.
        target: ``<provider>/default`` or ``<provider>/<model>``.
        version: The release its file comes from.
        summary: What matters about the target, in a sentence.
        instructions: What to do when writing a prompt for it.
        sources: Where each claim comes from.
    """

    category: str
    target: str
    version: str
    summary: str
    instructions: tuple[str, ...] = ()
    sources: tuple[str, ...] = ()

Part dataclass

One released file of the core rules that was read.

Attributes:

Name Type Description
category str

The category.

name str

The rule's folder name, or examples or guidance.

target str

default, <provider>/default or <provider>/<model>.

version str

The release it changed in.

released str

The day of that release.

changes tuple[str, ...]

What changed.

Source code in src/validia/rules/model.py
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
@dataclass(frozen=True, slots=True)
class Part:
    """One released file of the core rules that was read.

    Attributes:
        category: The category.
        name: The rule's folder name, or ``examples`` or ``guidance``.
        target: ``default``, ``<provider>/default`` or ``<provider>/<model>``.
        version: The release it changed in.
        released: The day of that release.
        changes: What changed.
    """

    category: str
    name: str
    target: str
    version: str
    released: str
    changes: tuple[str, ...]

Release dataclass

One release of one category: every file that changed in it, and why.

Attributes:

Name Type Description
category str

The category.

version str

The release, as in 1.1.0.

released str

The day of the release.

changes tuple[tuple[str, tuple[str, ...]], ...]

Each change, with the files that carry it, as show-reasoning anthropic/default.

Source code in src/validia/rules/model.py
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
@dataclass(frozen=True, slots=True)
class Release:
    """One release of one category: every file that changed in it, and why.

    Attributes:
        category: The category.
        version: The release, as in ``1.1.0``.
        released: The day of the release.
        changes: Each change, with the files that carry it, as ``show-reasoning anthropic/default``.
    """

    category: str
    version: str
    released: str
    changes: tuple[tuple[str, tuple[str, ...]], ...]

Rule dataclass

One check on prompt text.

Attributes:

Name Type Description
id str

Names the rule: its category and its folder, as in wording/capitals.

title str

What it finds, in a few words; the rule reference's heading.

pattern str

What to look for. With missing, what should be there.

fix str

What to do about a hit.

severity Severity

How much a hit matters, on the models in models.

scope tuple[Scope, ...]

The text it reads: the prompt, or tool descriptions.

case_sensitive bool

Whether case matters, as it does for capitals.

unless str

A hit is dropped when this matches in its sentence or the next.

run int

Fire once on this many consecutive matching sentences or lines.

unit Literal['sentence', 'line']

What run counts: sentence or line.

at_least int

Fire once when the pattern matches this many times.

missing bool

Fire when the pattern matches nowhere.

models tuple[str, ...]

Model names (globs) the severity applies to; empty means all.

otherwise Severity

The severity on every other model.

confidence Literal['high', 'med', 'low']

How sure a hit is: high, med or low.

judge str

The class set a judge would sort hits into, when one is needed.

category str

The category the rule belongs to.

fires tuple[str, ...]

Text the rule must fire on.

quiet tuple[str, ...]

Text the rule must stay quiet on.

origin tuple[str, ...]

Every file laid over it, in order, as anthropic/default 1.0.0 extend.

Source code in src/validia/rules/model.py
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
@dataclass(frozen=True, slots=True)
class Rule:
    """One check on prompt text.

    Attributes:
        id: Names the rule: its category and its folder, as in ``wording/capitals``.
        title: What it finds, in a few words; the rule reference's heading.
        pattern: What to look for. With ``missing``, what should be there.
        fix: What to do about a hit.
        severity: How much a hit matters, on the models in ``models``.
        scope: The text it reads: the prompt, or tool descriptions.
        case_sensitive: Whether case matters, as it does for capitals.
        unless: A hit is dropped when this matches in its sentence or the next.
        run: Fire once on this many consecutive matching sentences or lines.
        unit: What ``run`` counts: ``sentence`` or ``line``.
        at_least: Fire once when the pattern matches this many times.
        missing: Fire when the pattern matches nowhere.
        models: Model names (globs) the severity applies to; empty means all.
        otherwise: The severity on every other model.
        confidence: How sure a hit is: ``high``, ``med`` or ``low``.
        judge: The class set a judge would sort hits into, when one is needed.
        category: The category the rule belongs to.
        fires: Text the rule must fire on.
        quiet: Text the rule must stay quiet on.
        origin: Every file laid over it, in order, as anthropic/default 1.0.0 extend.
    """

    id: str
    pattern: str
    fix: str
    title: str = ""
    severity: Severity = "warn"
    scope: tuple[Scope, ...] = ("prompt",)
    case_sensitive: bool = False
    unless: str = ""
    run: int = 0
    unit: Literal["sentence", "line"] = "sentence"
    at_least: int = 0
    missing: bool = False
    models: tuple[str, ...] = ()
    otherwise: Severity = "info"
    confidence: Literal["high", "med", "low"] = "med"
    judge: str = ""
    category: str = ""
    fires: tuple[str, ...] = ()
    quiet: tuple[str, ...] = ()
    origin: tuple[str, ...] = ()

    def severity_for(self, model: str | None) -> Severity:
        """Say how much a hit matters on a model.

        Args:
            model: The model name, as in ``claude-opus-5-5``; ``None`` when unknown.

        Returns:
            ``severity`` on a listed model or when none are listed, else ``otherwise``.
        """
        if not self.models:
            return self.severity
        if model is not None and any(fnmatch.fnmatchcase(model, glob) for glob in self.models):
            return self.severity
        return self.otherwise

    def describe_severity(self, model: str | None) -> str:
        """Say how much a hit matters: on one model, or across the models it names.

        Args:
            model: The model name; ``None`` to describe every model the rule names.

        Returns:
            ``warn``, or ``error on claude-opus-5; warn elsewhere``.
        """
        if model is not None or not self.models:
            return self.severity_for(model)
        return f"{self.severity} on {', '.join(self.models)}; {self.otherwise} elsewhere"

    def describe_matching(self) -> str:
        """Say in words how this rule turns matches into findings.

        Returns:
            A sentence, as in ``Fires on every match.``
        """
        if self.missing:
            return "Fires when nothing matches: the pattern is what should be there."
        if self.run:
            unit = "sentences" if self.unit == "sentence" else "lines"
            return f"Fires once on {self.run} or more {unit} in a row that match."
        if self.at_least:
            return f"Fires once when the pattern matches {self.at_least} or more times."
        if self.unless:
            return "Fires on every match, unless its sentence or the next gives a reason."
        return "Fires on every match."

    def _flags(self) -> int:
        return re.MULTILINE | (0 if self.case_sensitive else re.IGNORECASE)

    def scan(self, text: str, model: str | None = None) -> list["Finding"]:
        """Find this rule's hits in a text.

        Args:
            text: The text to read.
            model: The target model, for the severity.

        Returns:
            The findings, in order.
        """
        pattern = re.compile(self.pattern, self._flags())
        severity = self.severity_for(model)
        if self.missing:
            if pattern.search(text):
                return []
            return [Finding(self.id, severity, 1, 1, "", self.fix)]
        if self.run:
            return self._runs(text, pattern, severity)
        matches = list(pattern.finditer(text))
        if self.at_least:
            if len(matches) < self.at_least:
                return []
            first = matches[0]
            return [
                _finding(self, severity, text, first.start(), f"{len(matches)} times: {first[0]}")
            ]
        if self.unless:
            reason = re.compile(self.unless, self._flags())
            sentences = _segments(text, "sentence")
            matches = [m for m in matches if not _explained(sentences, m.start(), reason)]
        return [_finding(self, severity, text, m.start(), m[0]) for m in matches]

    def _runs(self, text: str, pattern: re.Pattern[str], severity: Severity) -> list["Finding"]:
        """Fire once per run of ``run`` or more consecutive matching segments."""
        found: list[Finding] = []
        streak: list[tuple[int, str]] = []
        for start, segment in [*_segments(text, self.unit, blanks=True), (len(text), "")]:
            if segment and pattern.search(segment):
                streak.append((start, segment))
                continue
            if len(streak) >= self.run:
                first_start, first = streak[0]
                excerpt = f"{len(streak)} in a row: {first.strip()}"
                found.append(_finding(self, severity, text, first_start, excerpt))
            streak = []
        return found

severity_for

severity_for(model: str | None) -> Severity

Say how much a hit matters on a model.

Parameters:

Name Type Description Default
model str | None

The model name, as in claude-opus-5-5; None when unknown.

required

Returns:

Type Description
Severity

severity on a listed model or when none are listed, else otherwise.

Source code in src/validia/rules/model.py
169
170
171
172
173
174
175
176
177
178
179
180
181
182
def severity_for(self, model: str | None) -> Severity:
    """Say how much a hit matters on a model.

    Args:
        model: The model name, as in ``claude-opus-5-5``; ``None`` when unknown.

    Returns:
        ``severity`` on a listed model or when none are listed, else ``otherwise``.
    """
    if not self.models:
        return self.severity
    if model is not None and any(fnmatch.fnmatchcase(model, glob) for glob in self.models):
        return self.severity
    return self.otherwise

describe_severity

describe_severity(model: str | None) -> str

Say how much a hit matters: on one model, or across the models it names.

Parameters:

Name Type Description Default
model str | None

The model name; None to describe every model the rule names.

required

Returns:

Type Description
str

warn, or error on claude-opus-5; warn elsewhere.

Source code in src/validia/rules/model.py
184
185
186
187
188
189
190
191
192
193
194
195
def describe_severity(self, model: str | None) -> str:
    """Say how much a hit matters: on one model, or across the models it names.

    Args:
        model: The model name; ``None`` to describe every model the rule names.

    Returns:
        ``warn``, or ``error on claude-opus-5; warn elsewhere``.
    """
    if model is not None or not self.models:
        return self.severity_for(model)
    return f"{self.severity} on {', '.join(self.models)}; {self.otherwise} elsewhere"

describe_matching

describe_matching() -> str

Say in words how this rule turns matches into findings.

Returns:

Type Description
str

A sentence, as in Fires on every match.

Source code in src/validia/rules/model.py
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
def describe_matching(self) -> str:
    """Say in words how this rule turns matches into findings.

    Returns:
        A sentence, as in ``Fires on every match.``
    """
    if self.missing:
        return "Fires when nothing matches: the pattern is what should be there."
    if self.run:
        unit = "sentences" if self.unit == "sentence" else "lines"
        return f"Fires once on {self.run} or more {unit} in a row that match."
    if self.at_least:
        return f"Fires once when the pattern matches {self.at_least} or more times."
    if self.unless:
        return "Fires on every match, unless its sentence or the next gives a reason."
    return "Fires on every match."

scan

scan(text: str, model: str | None = None) -> list[Finding]

Find this rule's hits in a text.

Parameters:

Name Type Description Default
text str

The text to read.

required
model str | None

The target model, for the severity.

None

Returns:

Type Description
list[Finding]

The findings, in order.

Source code in src/validia/rules/model.py
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
def scan(self, text: str, model: str | None = None) -> list["Finding"]:
    """Find this rule's hits in a text.

    Args:
        text: The text to read.
        model: The target model, for the severity.

    Returns:
        The findings, in order.
    """
    pattern = re.compile(self.pattern, self._flags())
    severity = self.severity_for(model)
    if self.missing:
        if pattern.search(text):
            return []
        return [Finding(self.id, severity, 1, 1, "", self.fix)]
    if self.run:
        return self._runs(text, pattern, severity)
    matches = list(pattern.finditer(text))
    if self.at_least:
        if len(matches) < self.at_least:
            return []
        first = matches[0]
        return [
            _finding(self, severity, text, first.start(), f"{len(matches)} times: {first[0]}")
        ]
    if self.unless:
        reason = re.compile(self.unless, self._flags())
        sentences = _segments(text, "sentence")
        matches = [m for m in matches if not _explained(sentences, m.start(), reason)]
    return [_finding(self, severity, text, m.start(), m[0]) for m in matches]

RuleCheck dataclass

The outcome of proving one rule, or one example, against its cases.

Attributes:

Name Type Description
name str

The rule's id, or the example's name.

failures tuple[str, ...]

What went wrong; empty when it passed.

cases int

How many cases were run.

category str

The category it belongs to.

Source code in src/validia/rules/model.py
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
@dataclass(frozen=True, slots=True)
class RuleCheck:
    """The outcome of proving one rule, or one example, against its cases.

    Attributes:
        name: The rule's id, or the example's name.
        failures: What went wrong; empty when it passed.
        cases: How many cases were run.
        category: The category it belongs to.
    """

    name: str
    failures: tuple[str, ...]
    cases: int
    category: str = ""

    @property
    def passed(self) -> bool:
        """Whether every case behaved."""
        return not self.failures

passed property

passed: bool

Whether every case behaved.

RuleError

Bases: ValueError

Rules are malformed; problems lists every reason.

Attributes:

Name Type Description
source

The file or folder the problems are in.

problems

One line per problem.

Source code in src/validia/rules/model.py
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
class RuleError(ValueError):
    """Rules are malformed; ``problems`` lists every reason.

    Attributes:
        source: The file or folder the problems are in.
        problems: One line per problem.
    """

    def __init__(self, source: str, problems: Sequence[str]) -> None:
        """Keep the problems, and say them all in the message."""
        self.source = source
        self.problems = list(problems)
        count = len(self.problems)
        lines = "".join(f"\n  {problem}" for problem in self.problems)
        super().__init__(f"{source} has {count} problem{'s' if count != 1 else ''}:{lines}")

__init__

__init__(source: str, problems: Sequence[str]) -> None

Keep the problems, and say them all in the message.

Source code in src/validia/rules/model.py
112
113
114
115
116
117
118
def __init__(self, source: str, problems: Sequence[str]) -> None:
    """Keep the problems, and say them all in the message."""
    self.source = source
    self.problems = list(problems)
    count = len(self.problems)
    lines = "".join(f"\n  {problem}" for problem in self.problems)
    super().__init__(f"{source} has {count} problem{'s' if count != 1 else ''}:{lines}")

RulePack dataclass

A loaded, checked set of rules for one target.

Attributes:

Name Type Description
rules tuple[Rule, ...]

The rules, category by category.

examples tuple[Example, ...]

Whole-prompt examples that still hold for these rules.

parts tuple[Part, ...]

Every core file read.

guidance tuple[Guidance, ...]

What the vendor and model files say about writing for the target.

releases tuple[tuple[str, str], ...]

Each category, and the release it was read at.

target Target | None

The model the rules were resolved for, if any.

project str

The project folder laid over them, if any.

Source code in src/validia/rules/model.py
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
@dataclass(frozen=True, slots=True)
class RulePack:
    """A loaded, checked set of rules for one target.

    Attributes:
        rules: The rules, category by category.
        examples: Whole-prompt examples that still hold for these rules.
        parts: Every core file read.
        guidance: What the vendor and model files say about writing for the target.
        releases: Each category, and the release it was read at.
        target: The model the rules were resolved for, if any.
        project: The project folder laid over them, if any.
    """

    rules: tuple[Rule, ...]
    examples: tuple[Example, ...] = ()
    parts: tuple[Part, ...] = ()
    guidance: tuple[Guidance, ...] = ()
    releases: tuple[tuple[str, str], ...] = ()
    target: Target | None = None
    project: str = ""

    @property
    def version(self) -> str:
        """Name these rules as a report records them, so a result can be reproduced.

        Returns:
            The common release, then any category on another one, then the target and
            the project folder: ``1.0.0 (wording 1.1.0) for anthropic:claude-opus-5-5 + rules/``.
        """
        read = [(category, release) for category, release in self.releases if release]
        if not read:
            return f"{self.project}/" if self.project else "no rules"
        common = Counter(release for _, release in read).most_common(1)[0][0]
        odd = [f"{category} {release}" for category, release in read if release != common]
        named = f"{common} ({', '.join(odd)})" if odd else common
        if self.target is not None:
            named = f"{named} for {self.target[0]}:{self.target[1]}"
        return f"{named} + {self.project}/" if self.project else named

    @property
    def categories(self) -> tuple[str, ...]:
        """Every category in this pack, in order."""
        return tuple(category for category, _ in self.releases)

    @property
    def model(self) -> str | None:
        """The bare model name severities are read for."""
        return None if self.target is None else self.target[1]

    def only(self, categories: Sequence[str]) -> "RulePack":
        """Keep only some categories.

        Args:
            categories: The categories to keep.

        Returns:
            The narrower pack.
        """
        return RulePack(
            tuple(rule for rule in self.rules if rule.category in categories),
            tuple(example for example in self.examples if example.category in categories),
            tuple(part for part in self.parts if part.category in categories),
            tuple(note for note in self.guidance if note.category in categories),
            tuple(pair for pair in self.releases if pair[0] in categories),
            self.target,
            self.project,
        )

version property

version: str

Name these rules as a report records them, so a result can be reproduced.

Returns:

Type Description
str

The common release, then any category on another one, then the target and

str

the project folder: 1.0.0 (wording 1.1.0) for anthropic:claude-opus-5-5 + rules/.

categories property

categories: tuple[str, ...]

Every category in this pack, in order.

model property

model: str | None

The bare model name severities are read for.

only

only(categories: Sequence[str]) -> RulePack

Keep only some categories.

Parameters:

Name Type Description Default
categories Sequence[str]

The categories to keep.

required

Returns:

Type Description
RulePack

The narrower pack.

Source code in src/validia/rules/model.py
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
def only(self, categories: Sequence[str]) -> "RulePack":
    """Keep only some categories.

    Args:
        categories: The categories to keep.

    Returns:
        The narrower pack.
    """
    return RulePack(
        tuple(rule for rule in self.rules if rule.category in categories),
        tuple(example for example in self.examples if example.category in categories),
        tuple(part for part in self.parts if part.category in categories),
        tuple(note for note in self.guidance if note.category in categories),
        tuple(pair for pair in self.releases if pair[0] in categories),
        self.target,
        self.project,
    )

Written dataclass

What a write left in the project, and the proof that it holds.

Attributes:

Name Type Description
paths tuple[Path, ...]

The files written, the rule's first, then its cases if any.

pack RulePack

The project's rules for the target, as read after the write.

checks tuple[RuleCheck, ...]

The rule's own check, and every example that names it.

Source code in src/validia/rules/local.py
75
76
77
78
79
80
81
82
83
84
85
86
87
@dataclass(frozen=True, slots=True)
class Written:
    """What a write left in the project, and the proof that it holds.

    Attributes:
        paths: The files written, the rule's first, then its cases if any.
        pack: The project's rules for the target, as read after the write.
        checks: The rule's own check, and every example that names it.
    """

    paths: tuple[Path, ...]
    pack: RulePack
    checks: tuple[RuleCheck, ...]

RunSummary dataclass

A run, summed up.

Attributes:

Name Type Description
trials int

Trials run.

passed int

Trials that passed.

graded int

Trials that got a reply to grade.

errors Mapping[str, int]

Failure classes of the trials whose call failed, and how many of each.

interval tuple[float, float]

The 95% Wilson interval of the pass rate, as fractions.

cases tuple[CaseResult, ...]

Every case, in suite order.

groups Mapping[str, tuple[int, int]]

Each group's (passed, graded), in suite order.

input_tokens int

Tokens read, across the run.

output_tokens int

Tokens written, across the run.

cache_read_tokens int

Input tokens served from cache, across the run.

latency_p50_ms float

The median latency of the calls that answered.

latency_p95_ms float

Their 95th percentile.

served_models tuple[str, ...]

Every model the provider says answered.

Source code in src/validia/runs/runner.py
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
@dataclass(frozen=True, slots=True)
class RunSummary:
    """A run, summed up.

    Attributes:
        trials: Trials run.
        passed: Trials that passed.
        graded: Trials that got a reply to grade.
        errors: Failure classes of the trials whose call failed, and how many of each.
        interval: The 95% Wilson interval of the pass rate, as fractions.
        cases: Every case, in suite order.
        groups: Each group's ``(passed, graded)``, in suite order.
        input_tokens: Tokens read, across the run.
        output_tokens: Tokens written, across the run.
        cache_read_tokens: Input tokens served from cache, across the run.
        latency_p50_ms: The median latency of the calls that answered.
        latency_p95_ms: Their 95th percentile.
        served_models: Every model the provider says answered.
    """

    trials: int
    passed: int
    graded: int
    errors: Mapping[str, int]
    interval: tuple[float, float]
    cases: tuple[CaseResult, ...]
    groups: Mapping[str, tuple[int, int]]
    input_tokens: int
    output_tokens: int
    cache_read_tokens: int
    latency_p50_ms: float
    latency_p95_ms: float
    served_models: tuple[str, ...] = field(default=())

    @property
    def rate(self) -> float | None:
        """The pass rate over graded trials; ``None`` when nothing was graded."""
        return self.passed / self.graded if self.graded else None

    def to_dict(self) -> dict[str, Any]:
        """Render the summary as plain data, for ``summary.json``."""
        data = asdict(self)
        data["rate"] = self.rate
        data["groups"] = {name: {"passed": p, "graded": g} for name, (p, g) in self.groups.items()}
        return data

rate property

rate: float | None

The pass rate over graded trials; None when nothing was graded.

to_dict

to_dict() -> dict[str, Any]

Render the summary as plain data, for summary.json.

Source code in src/validia/runs/runner.py
446
447
448
449
450
451
def to_dict(self) -> dict[str, Any]:
    """Render the summary as plain data, for ``summary.json``."""
    data = asdict(self)
    data["rate"] = self.rate
    data["groups"] = {name: {"passed": p, "graded": g} for name, (p, g) in self.groups.items()}
    return data

Trial dataclass

One case, sent once.

Attributes:

Name Type Description
case Case

The case.

rep int

Which repetition this is, from 1.

Source code in src/validia/runs/runner.py
215
216
217
218
219
220
221
222
223
224
225
@dataclass(frozen=True, slots=True)
class Trial:
    """One case, sent once.

    Attributes:
        case: The case.
        rep: Which repetition this is, from 1.
    """

    case: Case
    rep: int

TrialResult dataclass

What one trial came to.

Attributes:

Name Type Description
case str

The case's id.

rep int

Which repetition, from 1.

tags tuple[str, ...]

The case's tags; the first is its group.

passed bool | None

Whether the reply was right; None when no reply came back.

reason str

Why it failed: the grader's reason, or the call's error. Empty on a pass.

reply str

The reply's text, for the record.

tool_calls tuple[str, ...]

The tools it called, as name(args).

error str | None

The failure class of a call that failed, as franca names it.

input_tokens int

Tokens the call read.

output_tokens int

Tokens it wrote.

cache_read_tokens int

Input tokens served from the provider's cache.

latency_ms float

How long the last attempt took on the wire.

attempts int

How many calls the trial took, retries included.

served_model str | None

The model the provider says answered.

stop_reason str | None

Why it stopped, as the provider says.

Source code in src/validia/runs/runner.py
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
@dataclass(frozen=True, slots=True)
class TrialResult:
    """What one trial came to.

    Attributes:
        case: The case's id.
        rep: Which repetition, from 1.
        tags: The case's tags; the first is its group.
        passed: Whether the reply was right; ``None`` when no reply came back.
        reason: Why it failed: the grader's reason, or the call's error. Empty on a pass.
        reply: The reply's text, for the record.
        tool_calls: The tools it called, as ``name(args)``.
        error: The failure class of a call that failed, as franca names it.
        input_tokens: Tokens the call read.
        output_tokens: Tokens it wrote.
        cache_read_tokens: Input tokens served from the provider's cache.
        latency_ms: How long the last attempt took on the wire.
        attempts: How many calls the trial took, retries included.
        served_model: The model the provider says answered.
        stop_reason: Why it stopped, as the provider says.
    """

    case: str
    rep: int
    tags: tuple[str, ...]
    passed: bool | None
    reason: str = ""
    reply: str = ""
    tool_calls: tuple[str, ...] = ()
    error: str | None = None
    input_tokens: int = 0
    output_tokens: int = 0
    cache_read_tokens: int = 0
    latency_ms: float = 0.0
    attempts: int = 1
    served_model: str | None = None
    stop_reason: str | None = None

    def to_json(self) -> str:
        """Render the result as one line of JSON, for ``trials.jsonl``."""
        return json.dumps(asdict(self), ensure_ascii=False)

to_json

to_json() -> str

Render the result as one line of JSON, for trials.jsonl.

Source code in src/validia/runs/runner.py
279
280
281
def to_json(self) -> str:
    """Render the result as one line of JSON, for ``trials.jsonl``."""
    return json.dumps(asdict(self), ensure_ascii=False)

BuildError

Bases: ValueError

The answers cannot make a prompt; problems lists every reason.

Attributes:

Name Type Description
problems

One line per problem, as in role: needs an answer.

Source code in src/validia/prompts/building.py
40
41
42
43
44
45
46
47
48
49
50
class BuildError(ValueError):
    """The answers cannot make a prompt; ``problems`` lists every reason.

    Attributes:
        problems: One line per problem, as in ``role: needs an answer``.
    """

    def __init__(self, problems: Sequence[str]) -> None:
        """Keep the problems, and say them all in the message."""
        self.problems = list(problems)
        super().__init__("; ".join(self.problems))

__init__

__init__(problems: Sequence[str]) -> None

Keep the problems, and say them all in the message.

Source code in src/validia/prompts/building.py
47
48
49
50
def __init__(self, problems: Sequence[str]) -> None:
    """Keep the problems, and say them all in the message."""
    self.problems = list(problems)
    super().__init__("; ".join(self.problems))

CaseSpec dataclass

A case to add: one input, and what a right answer looks like.

Attributes:

Name Type Description
id str

Unique within the suite; letters, digits, ., _ and -.

input str

The message the model gets.

expected Any

In the suite's answer type's shape: a label; a dict of JSON fields; a dict of text checks; or {"tool": ..., "args": {...}}.

tags tuple[str, ...]

Ordered labels; the first is the group results are reported by.

Source code in src/validia/suites/api.py
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
@dataclass(frozen=True, slots=True)
class CaseSpec:
    """A case to add: one input, and what a right answer looks like.

    Attributes:
        id: Unique within the suite; letters, digits, ``.``, ``_`` and ``-``.
        input: The message the model gets.
        expected: In the suite's answer type's shape: a label; a dict of JSON
            fields; a dict of text checks; or ``{"tool": ..., "args": {...}}``.
        tags: Ordered labels; the first is the group results are reported by.
    """

    id: str
    input: str
    expected: Any
    tags: tuple[str, ...] = ()

    @classmethod
    def from_dict(cls, data: Mapping[str, Any], where: str = "") -> "CaseSpec":
        """Build a case from a JSON object, as a request body carries it.

        Args:
            data: The object.
            where: A prefix for problem locations, such as ``cases[2].``.

        Returns:
            The case.

        Raises:
            SpecError: If a field is unknown, missing or of the wrong type.
        """
        problems = _fields(data, _CASE_FIELDS, where)
        problems += [
            f"{where}{key}: missing" for key in ("id", "input", "expected") if key not in data
        ]
        case = cls(
            id=_text(data, "id", where, problems),
            input=_text(data, "input", where, problems),
            expected=data.get("expected"),
            tags=_texts(data, "tags", where, problems),
        )
        if problems:
            raise SpecError(problems)
        return case

    def problems(self, taken: Collection[str] = (), where: str = "") -> list[str]:
        """Check what can be checked without the suite file.

        The expected answer's shape is checked against the suite itself when the
        case is written, by the same loader that reads every suite.

        Args:
            taken: Ids already in use.
            where: A prefix for problem locations.

        Returns:
            Every problem found.
        """
        found: list[str] = []
        if problem := name_problem(self.id):
            found.append(f"{where}id: {problem}")
        elif self.id in taken:
            found.append(f"{where}id: {self.id!r} is already a case")
        if not self.input.strip():
            found.append(f"{where}input: a case needs an input")
        if self.expected is None:
            found.append(f"{where}expected: missing")
        return found

from_dict classmethod

from_dict(
    data: Mapping[str, Any], where: str = ""
) -> CaseSpec

Build a case from a JSON object, as a request body carries it.

Parameters:

Name Type Description Default
data Mapping[str, Any]

The object.

required
where str

A prefix for problem locations, such as cases[2]..

''

Returns:

Type Description
CaseSpec

The case.

Raises:

Type Description
SpecError

If a field is unknown, missing or of the wrong type.

Source code in src/validia/suites/api.py
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
@classmethod
def from_dict(cls, data: Mapping[str, Any], where: str = "") -> "CaseSpec":
    """Build a case from a JSON object, as a request body carries it.

    Args:
        data: The object.
        where: A prefix for problem locations, such as ``cases[2].``.

    Returns:
        The case.

    Raises:
        SpecError: If a field is unknown, missing or of the wrong type.
    """
    problems = _fields(data, _CASE_FIELDS, where)
    problems += [
        f"{where}{key}: missing" for key in ("id", "input", "expected") if key not in data
    ]
    case = cls(
        id=_text(data, "id", where, problems),
        input=_text(data, "input", where, problems),
        expected=data.get("expected"),
        tags=_texts(data, "tags", where, problems),
    )
    if problems:
        raise SpecError(problems)
    return case

problems

problems(
    taken: Collection[str] = (), where: str = ""
) -> list[str]

Check what can be checked without the suite file.

The expected answer's shape is checked against the suite itself when the case is written, by the same loader that reads every suite.

Parameters:

Name Type Description Default
taken Collection[str]

Ids already in use.

()
where str

A prefix for problem locations.

''

Returns:

Type Description
list[str]

Every problem found.

Source code in src/validia/suites/api.py
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
def problems(self, taken: Collection[str] = (), where: str = "") -> list[str]:
    """Check what can be checked without the suite file.

    The expected answer's shape is checked against the suite itself when the
    case is written, by the same loader that reads every suite.

    Args:
        taken: Ids already in use.
        where: A prefix for problem locations.

    Returns:
        Every problem found.
    """
    found: list[str] = []
    if problem := name_problem(self.id):
        found.append(f"{where}id: {problem}")
    elif self.id in taken:
        found.append(f"{where}id: {self.id!r} is already a case")
    if not self.input.strip():
        found.append(f"{where}input: a case needs an input")
    if self.expected is None:
        found.append(f"{where}expected: missing")
    return found

Option dataclass

One answer a question offers.

Attributes:

Name Type Description
value str

The answer itself, as it is written into the prompt.

label str

A short description, when the value does not speak for itself.

Source code in src/validia/prompts/building.py
53
54
55
56
57
58
59
60
61
62
63
@dataclass(frozen=True, slots=True)
class Option:
    """One answer a question offers.

    Attributes:
        value: The answer itself, as it is written into the prompt.
        label: A short description, when the value does not speak for itself.
    """

    value: str
    label: str = ""

Question dataclass

One thing to ask, in a form any front end can show.

Attributes:

Name Type Description
id str

Where the answer goes: a section id, or :data:CLOSING.

text str

The question.

kind Literal['text', 'many', 'choice']

text takes one answer; many takes a list, one item per answer; choice offers :attr:options.

required bool

Whether an empty answer is refused, or skips the part.

example str

An answer that shows the expected shape.

options tuple[Option, ...]

For choice: the answers on offer.

custom bool

For choice: whether an answer outside the options is taken.

Source code in src/validia/prompts/building.py
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
@dataclass(frozen=True, slots=True)
class Question:
    """One thing to ask, in a form any front end can show.

    Attributes:
        id: Where the answer goes: a section id, or :data:`CLOSING`.
        text: The question.
        kind: ``text`` takes one answer; ``many`` takes a list, one item per
            answer; ``choice`` offers :attr:`options`.
        required: Whether an empty answer is refused, or skips the part.
        example: An answer that shows the expected shape.
        options: For ``choice``: the answers on offer.
        custom: For ``choice``: whether an answer outside the options is taken.
    """

    id: str
    text: str
    kind: Literal["text", "many", "choice"]
    required: bool = True
    example: str = ""
    options: tuple[Option, ...] = ()
    custom: bool = False

SpecError

Bases: ValueError

A request cannot be carried out; problems lists every reason.

Attributes:

Name Type Description
problems

One line per problem, as in cases[1].id: 'a' is used twice.

Source code in src/validia/suites/api.py
86
87
88
89
90
91
92
93
94
95
96
class SpecError(ValueError):
    """A request cannot be carried out; ``problems`` lists every reason.

    Attributes:
        problems: One line per problem, as in ``cases[1].id: 'a' is used twice``.
    """

    def __init__(self, problems: Sequence[str]) -> None:
        """Keep the problems, and say them all in the message."""
        self.problems = list(problems)
        super().__init__("; ".join(self.problems))

__init__

__init__(problems: Sequence[str]) -> None

Keep the problems, and say them all in the message.

Source code in src/validia/suites/api.py
93
94
95
96
def __init__(self, problems: Sequence[str]) -> None:
    """Keep the problems, and say them all in the message."""
    self.problems = list(problems)
    super().__init__("; ".join(self.problems))

SuiteSpec dataclass

A suite to create: where it goes, how it is graded, its prompt and its cases.

Attributes:

Name Type Description
name str

The suite's folder name.

answer AnswerType

What the prompt answers with: label, json, text, tool.

cases tuple[CaseSpec, ...]

At least one case.

folder str

The folder suites live in, relative to the project.

labels tuple[str, ...]

For label: at least two.

required tuple[str, ...]

For json: keys every reply must have.

prompt str

The prompt text, written to prompt.md; or use prompt_file.

prompt_file str

An existing prompt, relative to the project, that the suite points at instead of copying.

tools tuple[Tool, ...]

For tool: the tools, written to tools.json; or tools_file.

tools_file str

An existing tools file, relative to the project.

description str

What the suite measures.

Source code in src/validia/suites/api.py
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
@dataclass(frozen=True, slots=True)
class SuiteSpec:
    """A suite to create: where it goes, how it is graded, its prompt and its cases.

    Attributes:
        name: The suite's folder name.
        answer: What the prompt answers with: ``label``, ``json``, ``text``, ``tool``.
        cases: At least one case.
        folder: The folder suites live in, relative to the project.
        labels: For ``label``: at least two.
        required: For ``json``: keys every reply must have.
        prompt: The prompt text, written to ``prompt.md``; or use ``prompt_file``.
        prompt_file: An existing prompt, relative to the project, that the suite
            points at instead of copying.
        tools: For ``tool``: the tools, written to ``tools.json``; or ``tools_file``.
        tools_file: An existing tools file, relative to the project.
        description: What the suite measures.
    """

    name: str
    answer: AnswerType
    cases: tuple[CaseSpec, ...]
    folder: str = "evals"
    labels: tuple[str, ...] = ()
    required: tuple[str, ...] = ()
    prompt: str = ""
    prompt_file: str = ""
    tools: tuple[Tool, ...] = ()
    tools_file: str = ""
    description: str = ""

    @property
    def grade(self) -> Grade:
        """The grading this spec describes."""
        return Grade(self.answer, labels=self.labels, required=self.required)

    @classmethod
    def from_dict(cls, data: Mapping[str, Any]) -> "SuiteSpec":
        """Build a spec from a JSON object, as a request body carries it.

        Args:
            data: The object.

        Returns:
            The spec, its fields typed but not yet checked against each other;
            :func:`create_suite` does that.

        Raises:
            SpecError: If a field is unknown, missing or of the wrong type,
                listing every problem at once.
        """
        problems = _fields(data, _SUITE_FIELDS, "")
        problems += [f"{key}: missing" for key in ("name", "answer", "cases") if key not in data]
        answer = data.get("answer", "label")
        if answer not in ANSWER_TYPES:
            problems.append(f"answer: {answer!r} is not one of {list(ANSWER_TYPES)}")
            answer = "label"
        raw_cases = data.get("cases", [])
        cases: list[CaseSpec] = []
        if not isinstance(raw_cases, list):
            problems.append("cases: expected a list")
        else:
            for index, item in enumerate(raw_cases):
                if not isinstance(item, Mapping):
                    problems.append(f"cases[{index}]: expected an object")
                    continue
                try:
                    cases.append(CaseSpec.from_dict(item, f"cases[{index}]."))
                except SpecError as exc:
                    problems += exc.problems
        raw_tools = data.get("tools", [])
        tools: list[Tool] = []
        if not isinstance(raw_tools, list):
            problems.append("tools: expected a list")
        else:
            tools = [
                tool
                for index, item in enumerate(raw_tools)
                if (tool := _tool(item, f"tools[{index}].", problems)) is not None
            ]
        spec = cls(
            name=_text(data, "name", "", problems),
            answer=next(kind for kind in ANSWER_TYPES if kind == answer),
            cases=tuple(cases),
            folder=_text(data, "folder", "", problems) or "evals",
            labels=_texts(data, "labels", "", problems),
            required=_texts(data, "required", "", problems),
            prompt=_text(data, "prompt", "", problems),
            prompt_file=_text(data, "prompt_file", "", problems),
            tools=tuple(tools),
            tools_file=_text(data, "tools_file", "", problems),
            description=_text(data, "description", "", problems),
        )
        if problems:
            raise SpecError(problems)
        return spec

    def problems(self) -> list[str]:
        """Check the spec's parts against each other, without touching a disk.

        Returns:
            Every problem found.
        """
        found: list[str] = []
        if problem := name_problem(self.name):
            found.append(f"name: {problem}")
        if problem := folder_problem(self.folder):
            found.append(f"folder: {problem}")
        for key, value in (("prompt_file", self.prompt_file), ("tools_file", self.tools_file)):
            if value and (problem := folder_problem(value)):
                found.append(f"{key}: {problem.replace('a folder', 'a path')}")
        if self.answer == "label" and len(set(self.labels)) < 2:
            found.append("labels: a label suite needs at least two")
        if self.answer != "label" and self.labels:
            found.append("labels: only a label suite has labels")
        if self.answer != "json" and self.required:
            found.append("required: only a json suite has required keys")
        if self.answer == "tool" and bool(self.tools) == bool(self.tools_file):
            found.append("tools: a tool suite needs exactly one of tools and tools_file")
        if self.answer != "tool" and (self.tools or self.tools_file):
            found.append("tools: only a tool suite has tools")
        if bool(self.prompt.strip()) == bool(self.prompt_file):
            found.append("prompt: give exactly one of prompt and prompt_file")
        names = [tool.name for tool in self.tools]
        for index, tool in enumerate(self.tools):
            if tool.name == NO_TOOL:
                found.append(f"tools[{index}].name: {NO_TOOL!r} means no call at all")
            elif problem := name_problem(tool.name):
                found.append(f"tools[{index}].name: {problem}")
            elif names.count(tool.name) > 1 and names.index(tool.name) < index:
                found.append(f"tools[{index}].name: {tool.name!r} is defined twice")
        if not self.cases:
            found.append("cases: a suite needs at least one case")
        taken: set[str] = set()
        for index, case in enumerate(self.cases):
            found += case.problems(taken, f"cases[{index}].")
            taken.add(case.id)
        return found

grade property

grade: Grade

The grading this spec describes.

from_dict classmethod

from_dict(data: Mapping[str, Any]) -> SuiteSpec

Build a spec from a JSON object, as a request body carries it.

Parameters:

Name Type Description Default
data Mapping[str, Any]

The object.

required

Returns:

Type Description
SuiteSpec

The spec, its fields typed but not yet checked against each other;

SuiteSpec

func:create_suite does that.

Raises:

Type Description
SpecError

If a field is unknown, missing or of the wrong type, listing every problem at once.

Source code in src/validia/suites/api.py
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
@classmethod
def from_dict(cls, data: Mapping[str, Any]) -> "SuiteSpec":
    """Build a spec from a JSON object, as a request body carries it.

    Args:
        data: The object.

    Returns:
        The spec, its fields typed but not yet checked against each other;
        :func:`create_suite` does that.

    Raises:
        SpecError: If a field is unknown, missing or of the wrong type,
            listing every problem at once.
    """
    problems = _fields(data, _SUITE_FIELDS, "")
    problems += [f"{key}: missing" for key in ("name", "answer", "cases") if key not in data]
    answer = data.get("answer", "label")
    if answer not in ANSWER_TYPES:
        problems.append(f"answer: {answer!r} is not one of {list(ANSWER_TYPES)}")
        answer = "label"
    raw_cases = data.get("cases", [])
    cases: list[CaseSpec] = []
    if not isinstance(raw_cases, list):
        problems.append("cases: expected a list")
    else:
        for index, item in enumerate(raw_cases):
            if not isinstance(item, Mapping):
                problems.append(f"cases[{index}]: expected an object")
                continue
            try:
                cases.append(CaseSpec.from_dict(item, f"cases[{index}]."))
            except SpecError as exc:
                problems += exc.problems
    raw_tools = data.get("tools", [])
    tools: list[Tool] = []
    if not isinstance(raw_tools, list):
        problems.append("tools: expected a list")
    else:
        tools = [
            tool
            for index, item in enumerate(raw_tools)
            if (tool := _tool(item, f"tools[{index}].", problems)) is not None
        ]
    spec = cls(
        name=_text(data, "name", "", problems),
        answer=next(kind for kind in ANSWER_TYPES if kind == answer),
        cases=tuple(cases),
        folder=_text(data, "folder", "", problems) or "evals",
        labels=_texts(data, "labels", "", problems),
        required=_texts(data, "required", "", problems),
        prompt=_text(data, "prompt", "", problems),
        prompt_file=_text(data, "prompt_file", "", problems),
        tools=tuple(tools),
        tools_file=_text(data, "tools_file", "", problems),
        description=_text(data, "description", "", problems),
    )
    if problems:
        raise SpecError(problems)
    return spec

problems

problems() -> list[str]

Check the spec's parts against each other, without touching a disk.

Returns:

Type Description
list[str]

Every problem found.

Source code in src/validia/suites/api.py
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
def problems(self) -> list[str]:
    """Check the spec's parts against each other, without touching a disk.

    Returns:
        Every problem found.
    """
    found: list[str] = []
    if problem := name_problem(self.name):
        found.append(f"name: {problem}")
    if problem := folder_problem(self.folder):
        found.append(f"folder: {problem}")
    for key, value in (("prompt_file", self.prompt_file), ("tools_file", self.tools_file)):
        if value and (problem := folder_problem(value)):
            found.append(f"{key}: {problem.replace('a folder', 'a path')}")
    if self.answer == "label" and len(set(self.labels)) < 2:
        found.append("labels: a label suite needs at least two")
    if self.answer != "label" and self.labels:
        found.append("labels: only a label suite has labels")
    if self.answer != "json" and self.required:
        found.append("required: only a json suite has required keys")
    if self.answer == "tool" and bool(self.tools) == bool(self.tools_file):
        found.append("tools: a tool suite needs exactly one of tools and tools_file")
    if self.answer != "tool" and (self.tools or self.tools_file):
        found.append("tools: only a tool suite has tools")
    if bool(self.prompt.strip()) == bool(self.prompt_file):
        found.append("prompt: give exactly one of prompt and prompt_file")
    names = [tool.name for tool in self.tools]
    for index, tool in enumerate(self.tools):
        if tool.name == NO_TOOL:
            found.append(f"tools[{index}].name: {NO_TOOL!r} means no call at all")
        elif problem := name_problem(tool.name):
            found.append(f"tools[{index}].name: {problem}")
        elif names.count(tool.name) > 1 and names.index(tool.name) < index:
            found.append(f"tools[{index}].name: {tool.name!r} is defined twice")
    if not self.cases:
        found.append("cases: a suite needs at least one case")
    taken: set[str] = set()
    for index, case in enumerate(self.cases):
        found += case.problems(taken, f"cases[{index}].")
        taken.add(case.id)
    return found

Reply dataclass

Everything the model answered with, for one case.

Attributes:

Name Type Description
text str

The text of the answer.

tool_calls tuple[ToolCall, ...]

Tools it asked to call, in the order it asked.

Source code in src/validia/suites/expect.py
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
@dataclass(frozen=True, slots=True)
class Reply:
    """Everything the model answered with, for one case.

    Attributes:
        text: The text of the answer.
        tool_calls: Tools it asked to call, in the order it asked.
    """

    text: str = ""
    tool_calls: tuple[ToolCall, ...] = ()

    @property
    def empty(self) -> bool:
        """Whether the model said nothing and called nothing."""
        return not self.text.strip() and not self.tool_calls

empty property

empty: bool

Whether the model said nothing and called nothing.

ToolCall dataclass

One tool the model asked to call.

Attributes:

Name Type Description
name str

The tool's name.

args dict[str, Any]

Its arguments, decoded.

Source code in src/validia/suites/expect.py
37
38
39
40
41
42
43
44
45
46
47
@dataclass(frozen=True, slots=True)
class ToolCall:
    """One tool the model asked to call.

    Attributes:
        name: The tool's name.
        args: Its arguments, decoded.
    """

    name: str
    args: dict[str, Any] = field(default_factory=dict)

Verdict dataclass

The outcome of checking one reply.

Attributes:

Name Type Description
passed bool

Whether the reply is right.

reason str

Why it is not, when it is not; empty on a pass.

Source code in src/validia/suites/expect.py
68
69
70
71
72
73
74
75
76
77
78
@dataclass(frozen=True, slots=True)
class Verdict:
    """The outcome of checking one reply.

    Attributes:
        passed: Whether the reply is right.
        reason: Why it is not, when it is not; empty on a pass.
    """

    passed: bool
    reason: str = ""

Case dataclass

One input to run the prompt on, and what a right answer looks like.

Attributes:

Name Type Description
id str

Unique within the suite; results and transcripts are keyed by it.

input str

The user message sent to the model.

expected Expected

What the reply is checked against.

tags tuple[str, ...]

Ordered labels; the first is the group results are reported by.

Source code in src/validia/suites/suite.py
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
@dataclass(frozen=True, slots=True)
class Case:
    """One input to run the prompt on, and what a right answer looks like.

    Attributes:
        id: Unique within the suite; results and transcripts are keyed by it.
        input: The user message sent to the model.
        expected: What the reply is checked against.
        tags: Ordered labels; the first is the group results are reported by.
    """

    id: str
    input: str
    expected: Expected
    tags: tuple[str, ...] = ()

    def check(self, reply: Reply) -> Verdict:
        """Grade one reply to this case.

        An empty reply -- no text and no tool call -- always fails, whatever the
        answer type: no answer is not a wrong answer, and it must never score like
        a considered one.

        Args:
            reply: The model's answer.

        Returns:
            The verdict.
        """
        if reply.empty:
            return Verdict(passed=False, reason="empty reply")
        return self.expected.check(reply)

check

check(reply: Reply) -> Verdict

Grade one reply to this case.

An empty reply -- no text and no tool call -- always fails, whatever the answer type: no answer is not a wrong answer, and it must never score like a considered one.

Parameters:

Name Type Description Default
reply Reply

The model's answer.

required

Returns:

Type Description
Verdict

The verdict.

Source code in src/validia/suites/suite.py
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
def check(self, reply: Reply) -> Verdict:
    """Grade one reply to this case.

    An empty reply -- no text and no tool call -- always fails, whatever the
    answer type: no answer is not a wrong answer, and it must never score like
    a considered one.

    Args:
        reply: The model's answer.

    Returns:
        The verdict.
    """
    if reply.empty:
        return Verdict(passed=False, reason="empty reply")
    return self.expected.check(reply)

Grade dataclass

How each reply is scored.

Attributes:

Name Type Description
type AnswerType

The answer type, which decides the shape of every expected.

labels tuple[str, ...]

For label: the closed set of answers.

required tuple[str, ...]

For json: keys every reply must have.

Source code in src/validia/suites/suite.py
 98
 99
100
101
102
103
104
105
106
107
108
109
110
@dataclass(frozen=True, slots=True)
class Grade:
    """How each reply is scored.

    Attributes:
        type: The answer type, which decides the shape of every ``expected``.
        labels: For ``label``: the closed set of answers.
        required: For ``json``: keys every reply must have.
    """

    type: AnswerType
    labels: tuple[str, ...] = ()
    required: tuple[str, ...] = ()

Suite dataclass

A loaded, validated suite.

Attributes:

Name Type Description
path Path

The suite file.

prompt Path

The prompt file, resolved against the suite's directory.

grade Grade

How replies are scored.

cases tuple[Case, ...]

The cases, in file order.

description str

What the suite measures.

tools tuple[Tool, ...]

The tools the model is offered, if any.

tools_file Path | None

The file they were read from.

Source code in src/validia/suites/suite.py
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
@dataclass(frozen=True, slots=True)
class Suite:
    """A loaded, validated suite.

    Attributes:
        path: The suite file.
        prompt: The prompt file, resolved against the suite's directory.
        grade: How replies are scored.
        cases: The cases, in file order.
        description: What the suite measures.
        tools: The tools the model is offered, if any.
        tools_file: The file they were read from.
    """

    path: Path
    prompt: Path
    grade: Grade
    cases: tuple[Case, ...]
    description: str = ""
    tools: tuple[Tool, ...] = ()
    tools_file: Path | None = None

    def summary(self) -> str:
        """Describe the grader in one line, as ``run --dry-run`` prints it.

        Returns:
            The answer type, and what it checks against.
        """
        grade = self.grade
        if grade.type == "label":
            return f"label  ({', '.join(grade.labels)})"
        if grade.type == "tool":
            return f"tool  ({', '.join(tool.name for tool in self.tools)})"
        if grade.type == "json" and grade.required:
            return f"json  (required: {', '.join(grade.required)})"
        return grade.type

    def groups(self) -> dict[str, int]:
        """Count cases by their first tag, the group results are reported by.

        Returns:
            Case counts, in order of first appearance.
        """
        return dict(Counter(case.tags[0] if case.tags else UNTAGGED for case in self.cases))

summary

summary() -> str

Describe the grader in one line, as run --dry-run prints it.

Returns:

Type Description
str

The answer type, and what it checks against.

Source code in src/validia/suites/suite.py
186
187
188
189
190
191
192
193
194
195
196
197
198
199
def summary(self) -> str:
    """Describe the grader in one line, as ``run --dry-run`` prints it.

    Returns:
        The answer type, and what it checks against.
    """
    grade = self.grade
    if grade.type == "label":
        return f"label  ({', '.join(grade.labels)})"
    if grade.type == "tool":
        return f"tool  ({', '.join(tool.name for tool in self.tools)})"
    if grade.type == "json" and grade.required:
        return f"json  (required: {', '.join(grade.required)})"
    return grade.type

groups

groups() -> dict[str, int]

Count cases by their first tag, the group results are reported by.

Returns:

Type Description
dict[str, int]

Case counts, in order of first appearance.

Source code in src/validia/suites/suite.py
201
202
203
204
205
206
207
def groups(self) -> dict[str, int]:
    """Count cases by their first tag, the group results are reported by.

    Returns:
        Case counts, in order of first appearance.
    """
    return dict(Counter(case.tags[0] if case.tags else UNTAGGED for case in self.cases))

SuiteError

Bases: ValueError

A suite file is malformed; the message lists every problem in it.

Attributes:

Name Type Description
problems

One line per problem, for a caller that reports them as data.

Source code in src/validia/suites/suite.py
85
86
87
88
89
90
91
92
93
94
95
class SuiteError(ValueError):
    """A suite file is malformed; the message lists every problem in it.

    Attributes:
        problems: One line per problem, for a caller that reports them as data.
    """

    def __init__(self, message: str, problems: Sequence[str] = ()) -> None:
        """Keep the problems beside the rendered message."""
        super().__init__(message)
        self.problems = list(problems) or [message]

__init__

__init__(
    message: str, problems: Sequence[str] = ()
) -> None

Keep the problems beside the rendered message.

Source code in src/validia/suites/suite.py
92
93
94
95
def __init__(self, message: str, problems: Sequence[str] = ()) -> None:
    """Keep the problems beside the rendered message."""
    super().__init__(message)
    self.problems = list(problems) or [message]

Tool dataclass

A tool the model is offered, as franca's ToolDef takes it.

Attributes:

Name Type Description
name str

The tool's name, as the model will spell it in a call.

parameters dict[str, Any]

The JSON Schema of its arguments.

description str

What it does, in the model's own context window.

strict bool | None

Whether the provider must enforce the schema exactly.

Source code in src/validia/suites/suite.py
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
@dataclass(frozen=True, slots=True)
class Tool:
    """A tool the model is offered, as franca's ``ToolDef`` takes it.

    Attributes:
        name: The tool's name, as the model will spell it in a call.
        parameters: The JSON Schema of its arguments.
        description: What it does, in the model's own context window.
        strict: Whether the provider must enforce the schema exactly.
    """

    name: str
    parameters: dict[str, Any]
    description: str = ""
    strict: bool | None = None

default_template

default_template() -> PromptTemplate

Load the built-in template.

Returns:

Type Description
PromptTemplate

The template.

Source code in src/validia/prompts/template.py
378
379
380
381
382
383
384
def default_template() -> PromptTemplate:
    """Load the built-in template.

    Returns:
        The template.
    """
    return _parse(default_text(), "the built-in prompt template")

load_template

load_template(path: Path) -> PromptTemplate

Load a project's template.

Parameters:

Name Type Description Default
path Path

The prompt.toml file.

required

Returns:

Type Description
PromptTemplate

The template.

Raises:

Type Description
TemplateError

If it is malformed, listing every problem.

OSError

If it cannot be read.

Source code in src/validia/prompts/template.py
387
388
389
390
391
392
393
394
395
396
397
398
399
400
def load_template(path: Path) -> PromptTemplate:
    """Load a project's template.

    Args:
        path: The ``prompt.toml`` file.

    Returns:
        The template.

    Raises:
        TemplateError: If it is malformed, listing every problem.
        OSError: If it cannot be read.
    """
    return _parse(path.read_text(encoding="utf-8"), path.name)

core_categories

core_categories(
    folder: Traversable | Path | None = None,
) -> tuple[str, ...]

List the core categories, in reading order, then any others by name.

Parameters:

Name Type Description Default
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
tuple[str, ...]

The category names.

Source code in src/validia/rules/catalog.py
63
64
65
66
67
68
69
70
71
72
73
74
def core_categories(folder: Traversable | Path | None = None) -> tuple[str, ...]:
    """List the core categories, in reading order, then any others by name.

    Args:
        folder: Where the rules live; the package's own when omitted.

    Returns:
        The category names.
    """
    found = set(_dirs(_root(folder)))
    known = [category for category in CATEGORIES if category in found]
    return (*known, *sorted(found - set(known)))

core_releases

core_releases(
    category: str, folder: Traversable | Path | None = None
) -> tuple[Release, ...]

List a category's releases, oldest first, with what changed in each.

Parameters:

Name Type Description Default
category str

The category.

required
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
tuple[Release, ...]

The releases.

Raises:

Type Description
RuleError

If a file's header is malformed, or a release's files disagree on the day it was released.

Source code in src/validia/rules/catalog.py
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
def core_releases(category: str, folder: Traversable | Path | None = None) -> tuple[Release, ...]:
    """List a category's releases, oldest first, with what changed in each.

    Args:
        category: The category.
        folder: Where the rules live; the package's own when omitted.

    Returns:
        The releases.

    Raises:
        RuleError: If a file's header is malformed, or a release's files disagree on
            the day it was released.
    """
    base = _root(folder).joinpath(category)
    found: dict[str, list[tuple[str, str, tuple[str, ...]]]] = {}
    for name in _subjects(category, folder):
        for target in _targets_in(base.joinpath(name)):
            place = _at(base.joinpath(name), target)
            for version in _versions_in(place):
                source = f"{category}/{name}/{target}/{version}.toml"
                reader = _Reader()
                released, changes = _header(
                    reader, _load_toml(place.joinpath(f"{version}.toml"), source)
                )
                if reader.problems:
                    raise RuleError(source, reader.problems)
                found.setdefault(version, []).append((f"{name} {target}", released, changes))
    releases: list[Release] = []
    for version in _known_releases(category, folder):
        files = found[version]
        days = sorted({released for _, released, _ in files})
        if len(days) > 1:
            raise RuleError(
                f"{category} {version}", [f"its files disagree on the day: {', '.join(days)}"]
            )
        grouped: dict[str, list[str]] = {}
        for where, _, changes in files:
            for change in changes:
                grouped.setdefault(change, []).append(where)
        changed = tuple((change, tuple(wheres)) for change, wheres in grouped.items())
        releases.append(Release(category, version, days[0], changed))
    return tuple(releases)

core_rules

core_rules(
    category: str, folder: Traversable | Path | None = None
) -> tuple[str, ...]

List a category's rule folders, by name.

Parameters:

Name Type Description Default
category str

The category.

required
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
tuple[str, ...]

Rule names, such as show-reasoning; the rule's id is reasoning/show-reasoning.

Source code in src/validia/rules/catalog.py
77
78
79
80
81
82
83
84
85
86
87
def core_rules(category: str, folder: Traversable | Path | None = None) -> tuple[str, ...]:
    """List a category's rule folders, by name.

    Args:
        category: The category.
        folder: Where the rules live; the package's own when omitted.

    Returns:
        Rule names, such as ``show-reasoning``; the rule's id is ``reasoning/show-reasoning``.
    """
    return tuple(name for name in _dirs(_root(folder).joinpath(category)) if name not in _RESERVED)

core_targets

core_targets(
    category: str, folder: Traversable | Path | None = None
) -> tuple[str, ...]

List every target any file of a category is for: the default, then by vendor.

Parameters:

Name Type Description Default
category str

The category.

required
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
tuple[str, ...]

Targets such as anthropic/claude-opus-5-5.

Source code in src/validia/rules/catalog.py
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
def core_targets(category: str, folder: Traversable | Path | None = None) -> tuple[str, ...]:
    """List every target any file of a category is for: the default, then by vendor.

    Args:
        category: The category.
        folder: Where the rules live; the package's own when omitted.

    Returns:
        Targets such as ``anthropic/claude-opus-5-5``.
    """
    base = _root(folder).joinpath(category)
    found = {
        target
        for name in _subjects(category, folder)
        for target in _targets_in(base.joinpath(name))
    }
    return _ordered(found)

core_versions

core_versions(
    category: str,
    name: str,
    target: str = "default",
    folder: Traversable | Path | None = None,
) -> tuple[str, ...]

List the releases one rule, or a category's examples or guidance, changed in.

Parameters:

Name Type Description Default
category str

The category.

required
name str

The rule's folder name, or examples or guidance.

required
target str

The target folder.

'default'
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
tuple[str, ...]

Versions, oldest first; empty when there is no such folder.

Source code in src/validia/rules/catalog.py
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
def core_versions(
    category: str,
    name: str,
    target: str = "default",
    folder: Traversable | Path | None = None,
) -> tuple[str, ...]:
    """List the releases one rule, or a category's examples or guidance, changed in.

    Args:
        category: The category.
        name: The rule's folder name, or ``examples`` or ``guidance``.
        target: The target folder.
        folder: Where the rules live; the package's own when omitted.

    Returns:
        Versions, oldest first; empty when there is no such folder.
    """
    return tuple(_versions_in(_at(_root(folder).joinpath(category, name), target)))

default_rules

default_rules(
    target: Target | None = None,
    *,
    pins: Mapping[str, str] | None = None,
    categories: Sequence[str] | None = None,
    folder: Traversable | Path | None = None,
) -> RulePack

Load the core rules for a target: every rule, read along its chain.

Parameters:

Name Type Description Default
target Target | None

The model, as (provider, model); None reads only the default files, for every model alike.

None
pins Mapping[str, str] | None

Category names to the release to read; latest for the rest.

None
categories Sequence[str] | None

The categories to keep; every one when omitted.

None
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
RulePack

The combined pack.

Raises:

Type Description
RuleError

If a pin matches no release, or a file is malformed.

Source code in src/validia/rules/layering.py
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
def default_rules(
    target: Target | None = None,
    *,
    pins: Mapping[str, str] | None = None,
    categories: Sequence[str] | None = None,
    folder: Traversable | Path | None = None,
) -> RulePack:
    """Load the core rules for a target: every rule, read along its chain.

    Args:
        target: The model, as ``(provider, model)``; ``None`` reads only the default
            files, for every model alike.
        pins: Category names to the release to read; ``latest`` for the rest.
        categories: The categories to keep; every one when omitted.
        folder: Where the rules live; the package's own when omitted.

    Returns:
        The combined pack.

    Raises:
        RuleError: If a pin matches no release, or a file is malformed.
    """
    return _narrow(_pack(_core(target, pins, folder), target), categories)

lint_text

lint_text(
    text: str,
    pack: RulePack,
    *,
    model: str | None = None,
    scope: Scope = "prompt",
) -> list[Finding]

Run every rule for a kind of text over it.

Parameters:

Name Type Description Default
text str

The prompt, or one tool's description.

required
pack RulePack

The rules.

required
model str | None

The model name severities are read for; the pack's own when omitted.

None
scope Scope

What kind of text it is.

'prompt'

Returns:

Type Description
list[Finding]

Every finding, ordered by position.

Source code in src/validia/rules/matching.py
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
def lint_text(
    text: str,
    pack: RulePack,
    *,
    model: str | None = None,
    scope: Scope = "prompt",
) -> list[Finding]:
    """Run every rule for a kind of text over it.

    Args:
        text: The prompt, or one tool's description.
        pack: The rules.
        model: The model name severities are read for; the pack's own when omitted.
        scope: What kind of text it is.

    Returns:
        Every finding, ordered by position.
    """
    name = pack.model if model is None else model
    found = [
        finding for rule in pack.rules if scope in rule.scope for finding in rule.scan(text, name)
    ]
    return sorted(found, key=lambda finding: (finding.line, finding.column, finding.rule))

load_rules

load_rules(
    path: Path,
    *,
    target: Target | None = None,
    pins: Mapping[str, str] | None = None,
    categories: Sequence[str] | None = None,
    folder: Traversable | Path | None = None,
) -> RulePack

Load a project's rules/ folder, laid over the core rules for a target.

The folder has the core's layout and is read by the same code, after the core: each rule's chain runs through the core's default, vendor and model files, then the project's. Every file in the folder is checked, whichever categories are kept, and every problem is reported at once.

Parameters:

Name Type Description Default
path Path

The project's rules/ folder.

required
target Target | None

The model, as (provider, model).

None
pins Mapping[str, str] | None

Category names to the core release to read; latest for the rest.

None
categories Sequence[str] | None

The categories to keep, the project's own included.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
RulePack

The combined pack.

Raises:

Type Description
RuleError

If either side is malformed, listing every problem in the folder.

OSError

If a file cannot be read.

Source code in src/validia/rules/layering.py
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
def load_rules(
    path: Path,
    *,
    target: Target | None = None,
    pins: Mapping[str, str] | None = None,
    categories: Sequence[str] | None = None,
    folder: Traversable | Path | None = None,
) -> RulePack:
    """Load a project's ``rules/`` folder, laid over the core rules for a target.

    The folder has the core's layout and is read by the same code, after the core:
    each rule's chain runs through the core's default, vendor and model files, then the
    project's. Every file in the folder is checked, whichever categories are kept, and
    every problem is reported at once.

    Args:
        path: The project's ``rules/`` folder.
        target: The model, as ``(provider, model)``.
        pins: Category names to the core release to read; ``latest`` for the rest.
        categories: The categories to keep, the project's own included.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        The combined pack.

    Raises:
        RuleError: If either side is malformed, listing every problem in the folder.
        OSError: If a file cannot be read.
    """
    states = {state.name: state for state in _core(target, pins, folder)}
    owner = {
        f"{category}/{name}": category
        for category in core_categories(folder)
        for name in core_rules(category, folder)
    }
    chain = _chain(target)
    core = _Tree(_root(folder))

    def fetch(reader: _Reader, rule_id: str, raw: object) -> Rule | None:
        category = owner.get(rule_id)
        pin = str(raw)
        if category is None:
            reader.problems.append(
                f"from: {rule_id} is not a core rule, so it has no older release"
            )
            return None
        if not is_pin(pin) or pin == "latest":
            reader.problems.append('from: expected a release, as in "1" or "1.0.0"')
            return None
        try:
            older = _Category(category, _release_of(category, pin, folder))
        except RuleError as exc:
            reader.problems.append(f"from: {exc.problems[0]}")
            return None
        _read_rule(older, rule_id.split("/", 1)[1], chain, core, older.release)
        found = older.rules.get(rule_id)
        if found is None:
            reader.problems.append(f"from: {category} {older.release} has no rule {rule_id}")
        return found

    if not path.is_dir():
        raise RuleError(
            path.name, [f"expected a folder: {path.name}/<category>/<rule>/default/1.0.0.toml"]
        )
    tree = _Tree(path, f"{path.name}/", _PROJECT_RANK, core=False, fetch=fetch)
    problems: list[str] = []
    for entry in sorted(path.iterdir()):
        shown = f"{path.name}/{entry.name}"
        if entry.name.startswith((".", "_")):
            continue
        if not entry.is_dir():
            if entry.suffix == ".toml":
                problems.append(
                    f"{shown}: rules go in {path.name}/<category>/<rule>/default/1.0.0.toml"
                )
            continue
        state = states.setdefault(entry.name, _Category(entry.name, ""))
        for item in sorted(entry.iterdir()):
            where = f"{shown}/{item.name}"
            if item.name.startswith((".", "_")):
                continue
            if not item.is_dir():
                if item.suffix == ".toml":
                    problems.append(f"{where}: a rule's files go in {shown}/<rule>/default/")
                continue
            problems += _strays(item, where)
            if item.name in _RESERVED:
                continue
            try:
                _read_rule(state, item.name, chain, tree, "latest")
            except RuleError as exc:
                problems += [f"{exc.source}: {problem}" for problem in exc.problems]
                rule_id = f"{entry.name}/{item.name}"
                elsewhere = [other for other in owner if other.split("/", 1)[1] == item.name]
                if rule_id not in owner and elsewhere:
                    problems.append(
                        f"{where}: did you mean {elsewhere[0]}? Its folder is"
                        f" {path.name}/{elsewhere[0]}/"
                    )
        try:
            _read_extras(state, chain, tree, "latest")
        except RuleError as exc:
            problems += [f"{exc.source}: {problem}" for problem in exc.problems]
    if problems:
        raise RuleError(f"{path.name}/", problems)
    known = core_categories(folder)
    ordered = [states[name] for name in known] + [
        states[name] for name in sorted(states) if name not in known
    ]
    return _narrow(_pack(ordered, target, path.name), categories)

verify_core

verify_core(
    *,
    pins: Mapping[str, str] | None = None,
    categories: Sequence[str] | None = None,
    targets: Sequence[str] | None = None,
    folder: Traversable | Path | None = None,
) -> list[tuple[str, str, RulePack, list[RuleCheck]]]

Prove every core target in its own chain: each as it would be linted.

A vendor's files are proved over the default; a model's over both. Each examples file runs as its own model sees the rules, so a model's severities are tested where they are written.

Parameters:

Name Type Description Default
pins Mapping[str, str] | None

Category names to the release to read; latest for the rest.

None
categories Sequence[str] | None

The categories to prove; every core category when omitted.

None
targets Sequence[str] | None

The targets to prove, as anthropic/claude-opus-5-5; all when omitted.

None
folder Traversable | Path | None

Where the rules live; the package's own when omitted.

None

Returns:

Type Description
list[tuple[str, str, RulePack, list[RuleCheck]]]

(category, target, pack, checks) for every target proved.

Source code in src/validia/rules/layering.py
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
def verify_core(
    *,
    pins: Mapping[str, str] | None = None,
    categories: Sequence[str] | None = None,
    targets: Sequence[str] | None = None,
    folder: Traversable | Path | None = None,
) -> list[tuple[str, str, RulePack, list[RuleCheck]]]:
    """Prove every core target in its own chain: each as it would be linted.

    A vendor's files are proved over the default; a model's over both. Each examples
    file runs as its own model sees the rules, so a model's severities are tested where
    they are written.

    Args:
        pins: Category names to the release to read; ``latest`` for the rest.
        categories: The categories to prove; every core category when omitted.
        targets: The targets to prove, as ``anthropic/claude-opus-5-5``; all when omitted.
        folder: Where the rules live; the package's own when omitted.

    Returns:
        ``(category, target, pack, checks)`` for every target proved.
    """
    chosen = pins or {}
    out: list[tuple[str, str, RulePack, list[RuleCheck]]] = []
    for category in core_categories(folder) if categories is None else categories:
        for name in core_targets(category, folder):
            if targets is not None and name not in targets:
                continue
            provider, _, model = name.partition("/")
            if name == "default" or model == "default":
                target: Target | None = None
                chain = list(dict.fromkeys(["default", name]))
            else:
                target = (provider, model)
                chain = _chain(target)
            state = _read_category(category, chain, chosen.get(category, "latest"), folder)
            pack = _pack([state], target)
            out.append((category, name, pack, verify_rules(pack)))
    return out

verify_project

verify_project(
    path: Path,
    *,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> list[tuple[str, RulePack, list[RuleCheck]]]

Prove a project's rules/ folder for every target it has files for.

The default, then each vendor's and each model's, each read over the core, so a project's model-specific cases and examples are proved on their own model.

Parameters:

Name Type Description Default
path Path

The project's rules/ folder.

required
pins Mapping[str, str] | None

Category names to the core release to read; latest for the rest.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
list[tuple[str, RulePack, list[RuleCheck]]]

(target, pack, checks) for every target proved.

Raises:

Type Description
RuleError

If the folder or the core is malformed.

Source code in src/validia/rules/layering.py
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
def verify_project(
    path: Path,
    *,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> list[tuple[str, RulePack, list[RuleCheck]]]:
    """Prove a project's ``rules/`` folder for every target it has files for.

    The default, then each vendor's and each model's, each read over the core, so a
    project's model-specific cases and examples are proved on their own model.

    Args:
        path: The project's ``rules/`` folder.
        pins: Category names to the core release to read; ``latest`` for the rest.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        ``(target, pack, checks)`` for every target proved.

    Raises:
        RuleError: If the folder or the core is malformed.
    """
    found = {"default"}
    for category in (entry for entry in sorted(path.iterdir()) if entry.is_dir()):
        for item in (entry for entry in sorted(category.iterdir()) if entry.is_dir()):
            found |= set(_targets_in(item))
    out: list[tuple[str, RulePack, list[RuleCheck]]] = []
    for name in _ordered(found):
        provider, _, model = name.partition("/")
        target: Target | None = None if name == "default" else (provider, model)
        pack = load_rules(path, target=target, pins=pins, folder=folder)
        out.append((name, pack, verify_rules(pack)))
    return out

verify_rules

verify_rules(pack: RulePack) -> list[RuleCheck]

Prove every rule against its own cases, then every example.

An example is linted with its own category's rules, as its own model sees them, so a layer's examples prove the severities that layer sets.

Parameters:

Name Type Description Default
pack RulePack

The rules.

required

Returns:

Type Description
list[RuleCheck]

One check per rule, then one per example.

Source code in src/validia/rules/matching.py
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
def verify_rules(pack: RulePack) -> list[RuleCheck]:
    """Prove every rule against its own cases, then every example.

    An example is linted with its own category's rules, as its own model sees them, so
    a layer's examples prove the severities that layer sets.

    Args:
        pack: The rules.

    Returns:
        One check per rule, then one per example.
    """
    checks: list[RuleCheck] = []
    for rule in pack.rules:
        failures = [f"should fire on: {case!r}" for case in rule.fires if not rule.scan(case)]
        failures += [f"should stay quiet on: {case!r}" for case in rule.quiet if rule.scan(case)]
        cases = len(rule.fires) + len(rule.quiet)
        checks.append(RuleCheck(rule.id, tuple(failures), cases, rule.category))
    for example in pack.examples:
        scoped = pack.only([example.category]) if example.category else pack
        found = lint_text(example.text, scoped, model=example.model, scope=example.scope)
        fired = Counter(finding.rule for finding in found)
        failures = [f"{rule} should fire" for rule in sorted(example.fires - set(fired))]
        failures += [f"{rule} fired but should not" for rule in sorted(set(fired) - example.fires)]
        failures += [
            f"{rule} should fire {count} times, fired {fired[rule]}"
            for rule, count in sorted(example.counts.items())
            if fired[rule] != count
        ]
        for rule_id, severity in sorted(example.severities.items()):
            seen = sorted({finding.severity for finding in found if finding.rule == rule_id})
            if rule_id in fired and seen != [severity]:
                failures.append(f"{rule_id} should be {severity}, was {', '.join(seen)}")
        checks.append(RuleCheck(example.name, tuple(failures), 1, example.category))
    return checks

disable_rule

disable_rule(
    project: Path,
    rule_id: str,
    *,
    target: str = "default",
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written

Turn a rule off for a target and everything read after it.

Parameters:

Name Type Description Default
project Path

The project's rules/ folder; made when missing.

required
rule_id str

The rule, as reasoning/show-reasoning.

required
target str

default, <provider>/default or <provider>/<model>.

'default'
version str | None

The file's version; the first free one, 1.0.0, when omitted.

None
pins Mapping[str, str] | None

Category names to the core release to read.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
Written

What was written; it proves nothing of the rule, which is gone.

Raises:

Type Description
RuleError

If the rule is unknown, or the file exists; nothing is left written.

Source code in src/validia/rules/local.py
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
def disable_rule(
    project: Path,
    rule_id: str,
    *,
    target: str = "default",
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written:
    """Turn a rule off for a target and everything read after it.

    Args:
        project: The project's ``rules/`` folder; made when missing.
        rule_id: The rule, as ``reasoning/show-reasoning``.
        target: ``default``, ``<provider>/default`` or ``<provider>/<model>``.
        version: The file's version; the first free one, ``1.0.0``, when omitted.
        pins: Category names to the core release to read.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        What was written; it proves nothing of the rule, which is gone.

    Raises:
        RuleError: If the rule is unknown, or the file exists; nothing is left written.
    """
    _known(_pack(project, target, pins, folder), rule_id)
    rule = _rule_text(f"{rule_id} on {target}: turned off.", {"enabled": False})
    return _write(project, rule_id, target, rule, None, version, pins, folder)

extend_rule

extend_rule(
    project: Path,
    rule_id: str,
    *,
    target: str = "default",
    fields: Mapping[str, Any] | None = None,
    fires: Sequence[str] = (),
    quiet: Sequence[str] = (),
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written

Extend a rule for a target: change some fields, add cases, keep the rest.

Parameters:

Name Type Description Default
project Path

The project's rules/ folder; made when missing.

required
rule_id str

The rule, as reasoning/show-reasoning.

required
target str

default, <provider>/default or <provider>/<model>.

'default'
fields Mapping[str, Any] | None

The fields to change, as a rule file names them (severity, fix...).

None
fires Sequence[str]

More text the rule must fire on.

()
quiet Sequence[str]

More text the rule must stay quiet on.

()
version str | None

The file's version; the first free one, 1.0.0, when omitted. A newer version starts from the folder's newest file -- its fields and its cases -- and lays these changes over it, since it takes that file's place.

None
pins Mapping[str, str] | None

Category names to the core release to read.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
Written

What was written, and its proof.

Raises:

Type Description
RuleError

If there is nothing to change, the rule is unknown, the file exists, or the result does not load or prove; nothing is left written.

Source code in src/validia/rules/local.py
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
def extend_rule(
    project: Path,
    rule_id: str,
    *,
    target: str = "default",
    fields: Mapping[str, Any] | None = None,
    fires: Sequence[str] = (),
    quiet: Sequence[str] = (),
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written:
    """Extend a rule for a target: change some fields, add cases, keep the rest.

    Args:
        project: The project's ``rules/`` folder; made when missing.
        rule_id: The rule, as ``reasoning/show-reasoning``.
        target: ``default``, ``<provider>/default`` or ``<provider>/<model>``.
        fields: The fields to change, as a rule file names them (``severity``, ``fix``...).
        fires: More text the rule must fire on.
        quiet: More text the rule must stay quiet on.
        version: The file's version; the first free one, ``1.0.0``, when omitted. A newer
            version starts from the folder's newest file -- its fields and its cases --
            and lays these changes over it, since it takes that file's place.
        pins: Category names to the core release to read.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        What was written, and its proof.

    Raises:
        RuleError: If there is nothing to change, the rule is unknown, the file
            exists, or the result does not load or prove; nothing is left written.
    """
    changes = dict(fields or {})
    if not changes and not fires and not quiet:
        raise RuleError(
            "the extension", ["say what to change: a field, or cases it fires or stays quiet on"]
        )
    _known(_pack(project, target, pins, folder), rule_id)
    # A newer file takes the older one's place, so it starts from what the older one said.
    older, older_cases = _newest(project, rule_id, target) if version is not None else ({}, {})
    if "enabled" in older or "from" in older:
        raise RuleError(
            "the extension",
            [
                f"the newest file in that folder says {next(iter(older))} = ...; replace the rule instead"
            ],
        )
    body = {"extend": True, **older, **changes} if "pattern" not in older else {**older, **changes}
    fires = [*older_cases.get("fires", []), *fires]
    quiet = [*older_cases.get("quiet", []), *quiet]
    heading = "extended" if body.get("extend") else "replaces it"
    rule = _rule_text(f"{rule_id} on {target}: {heading}.", body)
    cases = _cases_text(rule_id, fires, quiet) if fires or quiet else None
    return _write(project, rule_id, target, rule, cases, version, pins, folder)

fall_back

fall_back(
    project: Path,
    rule_id: str,
    release: str,
    *,
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written

Take one core rule, along its whole chain, as an older release had it.

Parameters:

Name Type Description Default
project Path

The project's rules/ folder; made when missing.

required
rule_id str

The rule, as reasoning/show-reasoning.

required
release str

The core release, as 1 or 1.2.0.

required
version str | None

The file's version; the first free one, 1.0.0, when omitted.

None
pins Mapping[str, str] | None

Category names to the core release to read for everything else.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
Written

What was written, and its proof.

Raises:

Type Description
RuleError

If the release or rule is unknown, or the file exists; nothing is left written.

Source code in src/validia/rules/local.py
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
def fall_back(
    project: Path,
    rule_id: str,
    release: str,
    *,
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written:
    """Take one core rule, along its whole chain, as an older release had it.

    Args:
        project: The project's ``rules/`` folder; made when missing.
        rule_id: The rule, as ``reasoning/show-reasoning``.
        release: The core release, as ``1`` or ``1.2.0``.
        version: The file's version; the first free one, ``1.0.0``, when omitted.
        pins: Category names to the core release to read for everything else.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        What was written, and its proof.

    Raises:
        RuleError: If the release or rule is unknown, or the file exists; nothing is
            left written.
    """
    rule = _rule_text(f"{rule_id}: as core release {release} had it.", {"from": release})
    return _write(project, rule_id, "default", rule, None, version, pins, folder)

new_rule

new_rule(
    project: Path,
    rule_id: str,
    *,
    fields: Mapping[str, Any],
    fires: Sequence[str],
    quiet: Sequence[str],
    target: str = "default",
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written

Add a rule of the project's own, in a core category or one of its own.

Parameters:

Name Type Description Default
project Path

The project's rules/ folder; made when missing.

required
rule_id str

The new rule, as brand/sorry.

required
fields Mapping[str, Any]

Its fields; pattern and fix are required.

required
fires Sequence[str]

Text it must fire on; at least one.

required
quiet Sequence[str]

Text it must stay quiet on; at least one.

required
target str

default for every model, or a vendor's or model's folder for a rule only those models have.

'default'
version str | None

The file's version; the first free one, 1.0.0, when omitted.

None
pins Mapping[str, str] | None

Category names to the core release to read.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
Written

What was written, and its proof.

Raises:

Type Description
RuleError

If the rule exists already, a field is missing or wrong, or its cases fail; nothing is left written.

Source code in src/validia/rules/local.py
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
def new_rule(
    project: Path,
    rule_id: str,
    *,
    fields: Mapping[str, Any],
    fires: Sequence[str],
    quiet: Sequence[str],
    target: str = "default",
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written:
    """Add a rule of the project's own, in a core category or one of its own.

    Args:
        project: The project's ``rules/`` folder; made when missing.
        rule_id: The new rule, as ``brand/sorry``.
        fields: Its fields; ``pattern`` and ``fix`` are required.
        fires: Text it must fire on; at least one.
        quiet: Text it must stay quiet on; at least one.
        target: ``default`` for every model, or a vendor's or model's folder for a
            rule only those models have.
        version: The file's version; the first free one, ``1.0.0``, when omitted.
        pins: Category names to the core release to read.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        What was written, and its proof.

    Raises:
        RuleError: If the rule exists already, a field is missing or wrong, or its
            cases fail; nothing is left written.
    """
    _split(rule_id)
    if rule_id in {rule.id for rule in _pack(project, target, pins, folder).rules}:
        raise RuleError(
            "the rule", [f"{rule_id} exists already: extend it, or replace it, instead"]
        )
    ordered = {key: fields[key] for key in _FIELDS if key in fields}
    ordered.update({key: value for key, value in fields.items() if key not in ordered})
    rule = _rule_text(f"{rule_id}: a rule of this project's own.", ordered)
    cases = _cases_text(rule_id, fires, quiet)
    return _write(project, rule_id, target, rule, cases, version, pins, folder)

replace_rule

replace_rule(
    project: Path,
    rule_id: str,
    *,
    target: str = "default",
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written

Copy a rule, as the target reads it now, into a file of its own to edit.

The copy is a full definition, so it replaces the rule from the target up: every field that differs from the defaults, and every case the rule has.

Parameters:

Name Type Description Default
project Path

The project's rules/ folder; made when missing.

required
rule_id str

The rule, as reasoning/show-reasoning.

required
target str

default, <provider>/default or <provider>/<model>.

'default'
version str | None

The file's version; the first free one, 1.0.0, when omitted.

None
pins Mapping[str, str] | None

Category names to the core release to read.

None
folder Traversable | Path | None

Where the core rules live; the package's own when omitted.

None

Returns:

Type Description
Written

What was written, and its proof.

Raises:

Type Description
RuleError

If the rule is unknown, the file exists, or the copy does not load or prove; nothing is left written.

Source code in src/validia/rules/local.py
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
def replace_rule(
    project: Path,
    rule_id: str,
    *,
    target: str = "default",
    version: str | None = None,
    pins: Mapping[str, str] | None = None,
    folder: Traversable | Path | None = None,
) -> Written:
    """Copy a rule, as the target reads it now, into a file of its own to edit.

    The copy is a full definition, so it replaces the rule from the target up: every
    field that differs from the defaults, and every case the rule has.

    Args:
        project: The project's ``rules/`` folder; made when missing.
        rule_id: The rule, as ``reasoning/show-reasoning``.
        target: ``default``, ``<provider>/default`` or ``<provider>/<model>``.
        version: The file's version; the first free one, ``1.0.0``, when omitted.
        pins: Category names to the core release to read.
        folder: Where the core rules live; the package's own when omitted.

    Returns:
        What was written, and its proof.

    Raises:
        RuleError: If the rule is unknown, the file exists, or the copy does not
            load or prove; nothing is left written.
    """
    pack = _pack(project, target, pins, folder)
    _known(pack, rule_id)
    rule = next(item for item in pack.rules if item.id == rule_id)
    values = {
        "title": rule.title,
        "pattern": rule.pattern,
        "unless": rule.unless,
        "scope": list(rule.scope),
        "case_sensitive": rule.case_sensitive,
        "run": rule.run,
        "unit": rule.unit,
        "at_least": rule.at_least,
        "missing": rule.missing,
        "severity": rule.severity,
        "models": list(rule.models),
        "otherwise": rule.otherwise,
        "confidence": rule.confidence,
        "judge": rule.judge,
        "fix": rule.fix,
    }
    kept = {
        key: value
        for key, value in values.items()
        if key in ("pattern", "fix", "severity") or value != _DEFAULTS.get(key)
    }
    if not rule.models:
        kept.pop("otherwise", None)
    heading = f"{rule_id} on {target}: replaces it, copied from {' -> '.join(rule.origin)}."
    text = _rule_text(heading, kept)
    cases = _cases_text(rule_id, rule.fires, rule.quiet)
    return _write(project, rule_id, target, text, cases, version, pins, folder)

rule_target

rule_target(target: str) -> Target | None

Turn a target folder into the model it is read for.

Parameters:

Name Type Description Default
target str

default, <provider>/default or <provider>/<model>.

required

Returns:

Type Description
Target | None

None for the default; otherwise (provider, model), where a vendor's

Target | None

default reads as the model default of that vendor.

Raises:

Type Description
RuleError

If the target is not one of those shapes.

Source code in src/validia/rules/local.py
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
def rule_target(target: str) -> Target | None:
    """Turn a target folder into the model it is read for.

    Args:
        target: ``default``, ``<provider>/default`` or ``<provider>/<model>``.

    Returns:
        ``None`` for the default; otherwise ``(provider, model)``, where a vendor's
        default reads as the model ``default`` of that vendor.

    Raises:
        RuleError: If the target is not one of those shapes.
    """
    if target == "default":
        return None
    provider, slash, model = target.partition("/")
    if not slash or not provider or not model or "/" in model:
        raise RuleError(
            "the target",
            [f"{target!r}: expected default, <provider>/default or <provider>/<model>"],
        )
    return (provider, model)

build_model

build_model(
    provider: str,
    model: str,
    *,
    keys: KeyProvider,
    transport: Transport,
    clock: Clock,
    settings: ProviderSettings | None = None,
) -> ChatModel

Assemble franca's chat model for one provider and model id.

Parameters:

Name Type Description Default
provider str

The provider's slug, as in anthropic.

required
model str

The model id to send, as in claude-sonnet-5-5.

required
keys KeyProvider

Where the API key is read from, at call time.

required
transport Transport

The HTTP implementation; a scripted one in tests.

required
clock Clock

Source of time for latency and retry delays.

required
settings ProviderSettings | None

[providers.<name>]: a base_url (a gateway, a proxy), a timeout_s and extra_headers.

None

Returns:

Type Description
ChatModel

The model, ready for await model.complete(package).

Raises:

Type Description
AccessError

If franca has no chat endpoint for the provider.

Source code in src/validia/runs/runner.py
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
def build_model(
    provider: str,
    model: str,
    *,
    keys: KeyProvider,
    transport: Transport,
    clock: Clock,
    settings: ProviderSettings | None = None,
) -> ChatModel:
    """Assemble franca's chat model for one provider and model id.

    Args:
        provider: The provider's slug, as in ``anthropic``.
        model: The model id to send, as in ``claude-sonnet-5-5``.
        keys: Where the API key is read from, at call time.
        transport: The HTTP implementation; a scripted one in tests.
        clock: Source of time for latency and retry delays.
        settings: ``[providers.<name>]``: a ``base_url`` (a gateway, a proxy), a
            ``timeout_s`` and ``extra_headers``.

    Returns:
        The model, ready for ``await model.complete(package)``.

    Raises:
        AccessError: If franca has no chat endpoint for the provider.
    """
    route = _ROUTES.get(provider)
    if route is None:
        msg = f"cannot run {provider}:{model}: validia can call {', '.join(SUPPORTED)}"
        raise AccessError(msg)
    endpoint, adapter = route
    chosen = settings or ProviderSettings()
    if chosen.base_url:
        endpoint = endpoint.model_copy(update={"base_url": chosen.base_url})
    connector = Connector(
        endpoint,
        keys=keys,
        transport=transport,
        clock=clock,
        timeout_s=chosen.timeout_s,
        extra_headers=chosen.extra_headers,
    )
    return ChatModel(
        model=model,
        connector=connector,
        adapter=adapter(),
        profile=CHAT_PROFILES.resolve(Provider(provider), model),
        clock=clock,
    )

run_trial async

run_trial(
    client: ChatClient,
    suite: Suite,
    prompt: str,
    trial: Trial,
    *,
    clock: Clock,
    retries: int = 2,
) -> TrialResult

Send one trial, retrying what is worth retrying, and grade the reply.

A failure franca marks retryable -- a rate limit, an overloaded server, a dropped connection -- is retried after the delay the provider asked for, or after an exponential backoff when it asked for none. Anything else, and a retryable failure that outlives its retries, ends the trial as an error.

Parameters:

Name Type Description Default
client ChatClient

The model, or a pipeline of middleware around it.

required
suite Suite

The suite, for its tools.

required
prompt str

The prompt under test.

required
trial Trial

The trial.

required
clock Clock

Where retry delays are slept; injected, so tests never wait.

required
retries int

Extra attempts a retryable failure gets.

2

Returns:

Type Description
TrialResult

The trial's result. A failed call is a result, not an exception.

Source code in src/validia/runs/runner.py
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
async def run_trial(
    client: ChatClient,
    suite: Suite,
    prompt: str,
    trial: Trial,
    *,
    clock: Clock,
    retries: int = 2,
) -> TrialResult:
    """Send one trial, retrying what is worth retrying, and grade the reply.

    A failure franca marks retryable -- a rate limit, an overloaded server, a dropped
    connection -- is retried after the delay the provider asked for, or after an
    exponential backoff when it asked for none. Anything else, and a retryable failure
    that outlives its retries, ends the trial as an error.

    Args:
        client: The model, or a pipeline of middleware around it.
        suite: The suite, for its tools.
        prompt: The prompt under test.
        trial: The trial.
        clock: Where retry delays are slept; injected, so tests never wait.
        retries: Extra attempts a retryable failure gets.

    Returns:
        The trial's result. A failed call is a result, not an exception.
    """
    request = package(suite, prompt, trial.case)
    attempts = 0
    response: ModelResponse | None = None
    while response is None:
        attempts += 1
        try:
            response = await client.complete(request)
        except ModelError as exc:
            if not exc.retryable or attempts > retries:
                return TrialResult(
                    case=trial.case.id,
                    rep=trial.rep,
                    tags=trial.case.tags,
                    passed=None,
                    reason=exc.message,
                    error=str(exc.failure_class),
                    attempts=attempts,
                )
            backoff = min(_BACKOFF_CAP_S, _BACKOFF_S * 2 ** (attempts - 1))
            await clock.sleep(exc.retry_after_s if exc.retry_after_s is not None else backoff)
    reply = reply_of(response)
    verdict = trial.case.check(reply)
    usage = response.usage
    trace = response.trace
    return TrialResult(
        case=trial.case.id,
        rep=trial.rep,
        tags=trial.case.tags,
        passed=verdict.passed,
        reason=verdict.reason,
        reply=reply.text,
        tool_calls=tuple(
            f"{call.name}({json.dumps(call.args, sort_keys=True)})" for call in reply.tool_calls
        ),
        input_tokens=usage.input_tokens,
        output_tokens=usage.output_tokens,
        cache_read_tokens=usage.cache_read_tokens,
        latency_ms=round(trace.latency_ms, 1) if trace is not None else 0.0,
        attempts=attempts,
        served_model=response.served_model,
        stop_reason=response.stop_reason,
    )

summarize

summarize(results: Sequence[TrialResult]) -> RunSummary

Sum up a run's trial results.

Parameters:

Name Type Description Default
results Sequence[TrialResult]

Every trial's result, in trial order.

required

Returns:

Type Description
RunSummary

The summary: overall, by case, by group, and the cost.

Source code in src/validia/runs/runner.py
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
def summarize(results: Sequence[TrialResult]) -> RunSummary:
    """Sum up a run's trial results.

    Args:
        results: Every trial's result, in trial order.

    Returns:
        The summary: overall, by case, by group, and the cost.
    """
    graded = [result for result in results if result.passed is not None]
    passed = sum(1 for result in graded if result.passed)
    by_case: dict[str, list[TrialResult]] = {}
    for result in results:
        by_case.setdefault(result.case, []).append(result)
    cases = []
    groups: dict[str, list[int]] = {}
    for case, runs in by_case.items():
        group = runs[0].tags[0] if runs[0].tags else "untagged"
        ok = sum(1 for run in runs if run.passed)
        seen = sum(1 for run in runs if run.passed is not None)
        failure = next((run.reason for run in runs if not run.passed), "")
        cases.append(
            CaseResult(case, group, ok, seen, sum(1 for run in runs if run.passed is None), failure)
        )
        tally = groups.setdefault(group, [0, 0])
        tally[0] += ok
        tally[1] += seen
    latencies = [result.latency_ms for result in graded]
    return RunSummary(
        trials=len(results),
        passed=passed,
        graded=len(graded),
        errors=dict(Counter(result.error for result in results if result.error is not None)),
        interval=wilson(passed, len(graded)),
        cases=tuple(cases),
        groups={name: (tally[0], tally[1]) for name, tally in groups.items()},
        input_tokens=sum(result.input_tokens for result in results),
        output_tokens=sum(result.output_tokens for result in results),
        cache_read_tokens=sum(result.cache_read_tokens for result in results),
        latency_p50_ms=_percentile(latencies, 0.50),
        latency_p95_ms=_percentile(latencies, 0.95),
        served_models=tuple(
            dict.fromkeys(result.served_model for result in results if result.served_model)
        ),
    )

trials

trials(suite: Suite, reps: int) -> list[Trial]

Every trial a run makes: each case, reps times, case by case.

Parameters:

Name Type Description Default
suite Suite

The suite.

required
reps int

Repetitions of every case; at least 1.

required

Returns:

Type Description
list[Trial]

The trials, in the order they are reported.

Source code in src/validia/runs/runner.py
228
229
230
231
232
233
234
235
236
237
238
def trials(suite: Suite, reps: int) -> list[Trial]:
    """Every trial a run makes: each case, ``reps`` times, case by case.

    Args:
        suite: The suite.
        reps: Repetitions of every case; at least 1.

    Returns:
        The trials, in the order they are reported.
    """
    return [Trial(case, rep) for case in suite.cases for rep in range(1, reps + 1)]

unsupported

unsupported(suite: Suite) -> str | None

Say why a suite cannot be run through franca yet, if it cannot.

franca's adapters send text turns only, for now: a package's tools never reach the wire, and a tool call in the reply is not read back. A tool suite run anyway would grade every case as "called no tool" -- a wrong answer the model never gave -- so it is refused instead, until franca carries tools and validia's floor on it moves up.

Parameters:

Name Type Description Default
suite Suite

The suite.

required

Returns:

Type Description
str | None

The reason, or None when the suite can run.

Source code in src/validia/runs/runner.py
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
def unsupported(suite: Suite) -> str | None:
    """Say why a suite cannot be run through franca yet, if it cannot.

    franca's adapters send text turns only, for now: a package's tools never reach the
    wire, and a tool call in the reply is not read back. A tool suite run anyway would
    grade every case as "called no tool" -- a wrong answer the model never gave -- so it
    is refused instead, until franca carries tools and validia's floor on it moves up.

    Args:
        suite: The suite.

    Returns:
        The reason, or ``None`` when the suite can run.
    """
    if suite.tools or suite.grade.type == "tool":
        return (
            f"tool suites need tool calling, which franca {version('franca')} does not do yet:"
            " its adapters send text turns only, so the tools would never reach the model."
            " --dry-run still checks the suite"
        )
    return None

wilson

wilson(
    passed: int, total: int, z: float = 1.96
) -> tuple[float, float]

The Wilson score interval for a pass rate: honest at 0%, 100% and small n.

Parameters:

Name Type Description Default
passed int

Trials that passed.

required
total int

Trials graded.

required
z float

The normal quantile; 1.96 gives a 95% interval.

1.96

Returns:

Type Description
tuple[float, float]

(low, high) as fractions; (0.0, 1.0) when nothing was graded.

Source code in src/validia/runs/runner.py
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
def wilson(passed: int, total: int, z: float = 1.96) -> tuple[float, float]:
    """The Wilson score interval for a pass rate: honest at 0%, 100% and small n.

    Args:
        passed: Trials that passed.
        total: Trials graded.
        z: The normal quantile; 1.96 gives a 95% interval.

    Returns:
        ``(low, high)`` as fractions; ``(0.0, 1.0)`` when nothing was graded.
    """
    if total == 0:
        return 0.0, 1.0
    rate = passed / total
    centre = rate + z * z / (2 * total)
    spread = z * math.sqrt(rate * (1 - rate) / total + z * z / (4 * total * total))
    scale = 1 + z * z / total
    return max(0.0, (centre - spread) / scale), min(1.0, (centre + spread) / scale)

add_case

add_case(suite_path: Path, case: CaseSpec) -> Suite

Append one case to a suite, keeping it only if the suite still loads.

Parameters:

Name Type Description Default
suite_path Path

The suite file.

required
case CaseSpec

The case.

required

Returns:

Type Description
Suite

The suite as it now stands.

Raises:

Type Description
SpecError

If the case does not fit the suite, listing every problem.

Source code in src/validia/suites/api.py
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
def add_case(suite_path: Path, case: CaseSpec) -> Suite:
    """Append one case to a suite, keeping it only if the suite still loads.

    Args:
        suite_path: The suite file.
        case: The case.

    Returns:
        The suite as it now stands.

    Raises:
        SpecError: If the case does not fit the suite, listing every problem.
    """
    suite = load_suite(suite_path)
    problems = case.problems({existing.id for existing in suite.cases})
    if problems:
        raise SpecError(problems)
    try:
        return append_case(suite_path, case_toml(case.id, case.input, case.expected, case.tags))
    except (SuiteError, ValueError) as exc:
        raise SpecError(getattr(exc, "problems", [str(exc)])) from exc

check_reply

check_reply(
    suite: Suite, case_id: str, reply: Reply
) -> Verdict

Grade one reply against one case, without calling a model.

Parameters:

Name Type Description Default
suite Suite

The suite.

required
case_id str

The case's id.

required
reply Reply

The reply to grade.

required

Returns:

Type Description
Verdict

The verdict.

Raises:

Type Description
SpecError

If there is no such case, suggesting the closest one.

Source code in src/validia/suites/api.py
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
def check_reply(suite: Suite, case_id: str, reply: Reply) -> Verdict:
    """Grade one reply against one case, without calling a model.

    Args:
        suite: The suite.
        case_id: The case's id.
        reply: The reply to grade.

    Returns:
        The verdict.

    Raises:
        SpecError: If there is no such case, suggesting the closest one.
    """
    return find_case(suite, case_id).check(reply)

closing_options

closing_options(
    template: PromptTemplate,
    grade: Grade,
    tools: Sequence[Tool],
) -> list[str]

List the closing instructions on offer, filled in for this suite.

Parameters:

Name Type Description Default
template PromptTemplate

Where the instructions come from.

required
grade Grade

The suite's grading: its answer type, labels and keys.

required
tools Sequence[Tool]

The suite's tools.

required

Returns:

Type Description
list[str]

The instructions, the template's default first.

Source code in src/validia/prompts/building.py
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
def closing_options(template: PromptTemplate, grade: Grade, tools: Sequence[Tool]) -> list[str]:
    """List the closing instructions on offer, filled in for this suite.

    Args:
        template: Where the instructions come from.
        grade: The suite's grading: its answer type, labels and keys.
        tools: The suite's tools.

    Returns:
        The instructions, the template's default first.
    """
    return [
        fill_output(option, labels=grade.labels, keys=grade.required, tools=tools)
        for option in template.output[grade.type]
    ]

create_suite

create_suite(spec: SuiteSpec, root: Path) -> Suite

Create a suite: its file, its prompt and tools when given as text, all at once.

Nothing is written unless the whole suite loads; a failure leaves no trace.

Parameters:

Name Type Description Default
spec SuiteSpec

What to create.

required
root Path

The project the spec's folders and files are relative to.

required

Returns:

Type Description
Suite

The suite, as loaded from what was written.

Raises:

Type Description
SpecError

If the spec, a referenced file, or the written suite has a problem, listing every one.

Source code in src/validia/suites/api.py
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
def create_suite(spec: SuiteSpec, root: Path) -> Suite:
    """Create a suite: its file, its prompt and tools when given as text, all at once.

    Nothing is written unless the whole suite loads; a failure leaves no trace.

    Args:
        spec: What to create.
        root: The project the spec's folders and files are relative to.

    Returns:
        The suite, as loaded from what was written.

    Raises:
        SpecError: If the spec, a referenced file, or the written suite has a
            problem, listing every one.
    """
    problems = spec.problems()
    if problems:
        raise SpecError(problems)
    folder = root / spec.folder / spec.name
    if not _inside(folder, root):
        problems.append(f"folder: {spec.folder}/{spec.name} leads outside the project")
    elif folder.exists():
        problems.append(f"name: {spec.folder}/{spec.name} already exists")
    for key, value in (("prompt_file", spec.prompt_file), ("tools_file", spec.tools_file)):
        if value and not _inside(root / value, root):
            problems.append(f"{key}: {value!r} leads outside the project")
    if problems:
        raise SpecError(problems)
    prompt_file = root / spec.prompt_file if spec.prompt_file else None
    if prompt_file is not None and not prompt_file.is_file():
        problems.append(f"prompt_file: no such file {spec.prompt_file!r}")
    tools = spec.tools
    if spec.tools_file:
        try:
            tools = load_tools(root / spec.tools_file)
        except SuiteError as exc:
            problems += [f"tools_file: {problem}" for problem in exc.problems]
    if problems:
        raise SpecError(problems)

    files: dict[str, str] = {}
    prompt_ref = "prompt.md"
    if prompt_file is not None:
        prompt_ref = _relative(prompt_file, folder)
    else:
        files["prompt.md"] = spec.prompt if spec.prompt.endswith("\n") else f"{spec.prompt}\n"
    tools_ref = ""
    if spec.tools:
        definitions = [
            {"name": tool.name, "description": tool.description, "parameters": tool.parameters}
            for tool in tools
        ]
        files["tools.json"] = json.dumps(definitions, indent=2) + "\n"
        tools_ref = "tools.json"
    elif spec.tools_file:
        tools_ref = _relative(root / spec.tools_file, folder)

    blocks: list[str] = []
    for index, case in enumerate(spec.cases):
        try:
            blocks.append(case_toml(case.id, case.input, case.expected, case.tags))
        except ValueError as exc:
            problems.append(f"cases[{index}].expected: {exc}")
    if problems:
        raise SpecError(problems)

    shown = PurePosixPath(spec.folder) / spec.name / "suite.toml"
    rendered = suite_toml(
        grade=spec.grade,
        prompt=prompt_ref,
        cases=blocks,
        run=shown.as_posix(),
        description=spec.description,
        tools=tools_ref,
    )
    try:
        return write_suite(folder, {"suite.toml": rendered, **files})
    except SuiteError as exc:
        raise SpecError(exc.problems) from exc

describe_suite

describe_suite(
    suite: Suite, root: Path | None = None
) -> dict[str, Any]

Describe a suite as JSON-ready data: what a front end needs to show it.

Parameters:

Name Type Description Default
suite Suite

The suite.

required
root Path | None

Paths are shown relative to this, when they are inside it.

None

Returns:

Type Description
dict[str, Any]

The answer type, labels, required keys, tools with their parameters, and

dict[str, Any]

every case's id, tags and input.

Source code in src/validia/suites/api.py
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
def describe_suite(suite: Suite, root: Path | None = None) -> dict[str, Any]:
    """Describe a suite as JSON-ready data: what a front end needs to show it.

    Args:
        suite: The suite.
        root: Paths are shown relative to this, when they are inside it.

    Returns:
        The answer type, labels, required keys, tools with their parameters, and
        every case's id, tags and input.
    """

    def shown(path: Path) -> str:
        if root is not None and path.resolve().is_relative_to(root.resolve()):
            return path.resolve().relative_to(root.resolve()).as_posix()
        return path.as_posix()

    return {
        "path": shown(suite.path),
        "prompt": shown(suite.prompt),
        "description": suite.description,
        "answer": suite.grade.type,
        "labels": list(suite.grade.labels),
        "required": list(suite.grade.required),
        "tools": [
            {"name": tool.name, "description": tool.description, "parameters": tool.parameters}
            for tool in suite.tools
        ],
        "groups": suite.groups(),
        "cases": [{"id": c.id, "tags": list(c.tags), "input": c.input} for c in suite.cases],
    }

find_case

find_case(suite: Suite, case_id: str) -> Case

Look a case up by id.

Parameters:

Name Type Description Default
suite Suite

The suite.

required
case_id str

The case's id.

required

Returns:

Type Description
Case

The case.

Raises:

Type Description
SpecError

If there is no such case, suggesting the closest one.

Source code in src/validia/suites/api.py
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
def find_case(suite: Suite, case_id: str) -> Case:
    """Look a case up by id.

    Args:
        suite: The suite.
        case_id: The case's id.

    Returns:
        The case.

    Raises:
        SpecError: If there is no such case, suggesting the closest one.
    """
    cases = {case.id: case for case in suite.cases}
    if case_id not in cases:
        close = difflib.get_close_matches(case_id, list(cases), n=1)
        hint = f"; did you mean {close[0]!r}?" if close else ""
        raise SpecError([f"no case {case_id!r} in {suite.path.name}{hint}"])
    return cases[case_id]

locate_suite

locate_suite(
    root: Path, name: str, folder: str = "evals"
) -> Path

Find a suite file by name, refusing any name that would leave the project.

A front end that takes a suite's name from a URL or a form should come through here rather than joining paths itself: .. in a name is a request for a file the API was never meant to touch.

Parameters:

Name Type Description Default
root Path

The project.

required
name str

The suite's name.

required
folder str

The folder suites live in, relative to the project.

'evals'

Returns:

Type Description
Path

The suite file's path.

Raises:

Type Description
SpecError

If the name or folder is not a plain path inside the project, or there is no such suite.

Source code in src/validia/suites/api.py
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
def locate_suite(root: Path, name: str, folder: str = "evals") -> Path:
    """Find a suite file by name, refusing any name that would leave the project.

    A front end that takes a suite's name from a URL or a form should come through
    here rather than joining paths itself: ``..`` in a name is a request for a file
    the API was never meant to touch.

    Args:
        root: The project.
        name: The suite's name.
        folder: The folder suites live in, relative to the project.

    Returns:
        The suite file's path.

    Raises:
        SpecError: If the name or folder is not a plain path inside the project,
            or there is no such suite.
    """
    problems = [f"name: {problem}" for problem in [name_problem(name)] if problem]
    problems += [f"folder: {problem}" for problem in [folder_problem(folder)] if problem]
    path = root / folder / name / "suite.toml"
    if not problems and not _inside(path, root):
        problems.append(f"name: {folder}/{name} leads outside the project")
    if not problems and not path.is_file():
        problems.append(f"name: no suite {folder}/{name}")
    if problems:
        raise SpecError(problems)
    return path

parse_reply

parse_reply(data: Mapping[str, Any]) -> Reply

Build a reply from a JSON object: {"text": ..., "tool_calls": [...]}.

Parameters:

Name Type Description Default
data Mapping[str, Any]

The object; each tool call is {"name": ..., "args": {...}}.

required

Returns:

Type Description
Reply

The reply.

Raises:

Type Description
SpecError

If a field is unknown or of the wrong type.

Source code in src/validia/suites/api.py
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
def parse_reply(data: Mapping[str, Any]) -> Reply:
    """Build a reply from a JSON object: ``{"text": ..., "tool_calls": [...]}``.

    Args:
        data: The object; each tool call is ``{"name": ..., "args": {...}}``.

    Returns:
        The reply.

    Raises:
        SpecError: If a field is unknown or of the wrong type.
    """
    problems = _fields(data, ("text", "tool_calls"), "")
    text = _text(data, "text", "", problems)
    calls: list[ToolCall] = []
    raw = data.get("tool_calls", [])
    if not isinstance(raw, list):
        problems.append("tool_calls: expected a list")
        raw = []
    for index, call in enumerate(raw):
        where = f"tool_calls[{index}]."
        if not isinstance(call, Mapping) or not isinstance(call.get("name"), str):
            problems.append(f"{where}name: expected text")
            continue
        problems += _fields(call, ("name", "args"), where)
        args = call.get("args", {})
        if not isinstance(args, dict):
            problems.append(f"{where}args: expected an object")
            args = {}
        calls.append(ToolCall(call["name"], args))
    if problems:
        raise SpecError(problems)
    return Reply(text=text, tool_calls=tuple(calls))

prompt_questions

prompt_questions(
    template: PromptTemplate,
    grade: Grade,
    tools: Sequence[Tool] = (),
) -> tuple[Question, ...]

List the questions that build a prompt for this suite, in order.

Parameters:

Name Type Description Default
template PromptTemplate

The sections to ask about.

required
grade Grade

The suite's grading, which picks the sections and the closing.

required
tools Sequence[Tool]

The suite's tools, listed in a tool prompt's closing.

()

Returns:

Type Description
tuple[Question, ...]

One question per section that applies, then :data:CLOSING.

Source code in src/validia/prompts/building.py
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
def prompt_questions(
    template: PromptTemplate, grade: Grade, tools: Sequence[Tool] = ()
) -> tuple[Question, ...]:
    """List the questions that build a prompt for this suite, in order.

    Args:
        template: The sections to ask about.
        grade: The suite's grading, which picks the sections and the closing.
        tools: The suite's tools, listed in a tool prompt's closing.

    Returns:
        One question per section that applies, then :data:`CLOSING`.
    """
    questions: list[Question] = []
    for section in template.sections_for(grade.type):
        if section.choices:
            options = tuple(Option(choice) for choice in section.choices)
            question = Question(
                section.id,
                section.ask,
                "choice",
                section.required,
                section.example,
                options,
                custom=True,
            )
        else:
            kind: Literal["many", "text"] = "many" if section.many else "text"
            question = Question(section.id, section.ask, kind, section.required, section.example)
        questions.append(question)
    closings = tuple(Option(text) for text in closing_options(template, grade, tools))
    questions.append(
        Question(CLOSING, "How should it answer?", "choice", options=closings, custom=True)
    )
    return tuple(questions)

render_prompt

render_prompt(
    template: PromptTemplate,
    grade: Grade,
    answers: Mapping[str, Answer],
    tools: Sequence[Tool] = (),
) -> str

Turn a complete set of answers into the prompt.

Parameters:

Name Type Description Default
template PromptTemplate

The sections the answers fill.

required
grade Grade

The suite's grading, which picks the sections that apply.

required
answers Mapping[str, Answer]

Question ids, as :func:prompt_questions lists them, to answers. A skipped optional question may be left out or answered with "".

required
tools Sequence[Tool]

The suite's tools.

()

Returns:

Type Description
str

The prompt, sections separated by blank lines, ending in a newline.

Raises:

Type Description
BuildError

If an answer is missing, misshapen, or for no question, listing every problem at once.

Source code in src/validia/prompts/building.py
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
def render_prompt(
    template: PromptTemplate,
    grade: Grade,
    answers: Mapping[str, Answer],
    tools: Sequence[Tool] = (),
) -> str:
    """Turn a complete set of answers into the prompt.

    Args:
        template: The sections the answers fill.
        grade: The suite's grading, which picks the sections that apply.
        answers: Question ids, as :func:`prompt_questions` lists them, to answers.
            A skipped optional question may be left out or answered with ``""``.
        tools: The suite's tools.

    Returns:
        The prompt, sections separated by blank lines, ending in a newline.

    Raises:
        BuildError: If an answer is missing, misshapen, or for no question,
            listing every problem at once.
    """
    questions = prompt_questions(template, grade, tools)
    known = {question.id for question in questions}
    problems = [f"{key}: no such question" for key in answers if key not in known]
    sections = {section.id: section for section in template.sections_for(grade.type)}
    parts: list[str] = []
    for question in questions:
        empty: Answer = [] if question.kind == "many" else ""
        items = _items(question, answers.get(question.id, empty), problems)
        if question.required and not items:
            problems.append(f"{question.id}: needs an answer")
        if question.id == CLOSING:
            parts.extend(items[:1])
        elif items:
            parts.append(sections[question.id].render(items))
    if problems:
        raise BuildError(problems)
    return "\n\n".join(part for part in parts if part) + "\n"

review_answer

review_answer(
    template: PromptTemplate, question: str, answer: str
) -> tuple[Hint, ...]

Run the template's core checks on one answer.

A hint is advice, never a refusal: the front end shows it and the person decides whether to rewrite.

Parameters:

Name Type Description Default
template PromptTemplate

Where the checks come from.

required
question str

The id of the question the answer is for.

required
answer str

The answer.

required

Returns:

Type Description
tuple[Hint, ...]

The hints it sets off, in template order.

Source code in src/validia/prompts/building.py
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
def review_answer(template: PromptTemplate, question: str, answer: str) -> tuple[Hint, ...]:
    """Run the template's core checks on one answer.

    A hint is advice, never a refusal: the front end shows it and the person
    decides whether to rewrite.

    Args:
        template: Where the checks come from.
        question: The id of the question the answer is for.
        answer: The answer.

    Returns:
        The hints it sets off, in template order.
    """
    return tuple(template.hints_for(question, answer)) if answer else ()

review_answers

review_answers(
    template: PromptTemplate, answers: Mapping[str, Answer]
) -> dict[str, tuple[Hint, ...]]

Run the core checks on a whole set of answers, as a form submits them.

Parameters:

Name Type Description Default
template PromptTemplate

Where the checks come from.

required
answers Mapping[str, Answer]

Question ids to answers.

required

Returns:

Type Description
dict[str, tuple[Hint, ...]]

Question ids to the hints their answers set off; quiet ones left out.

Source code in src/validia/prompts/building.py
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
def review_answers(
    template: PromptTemplate, answers: Mapping[str, Answer]
) -> dict[str, tuple[Hint, ...]]:
    """Run the core checks on a whole set of answers, as a form submits them.

    Args:
        template: Where the checks come from.
        answers: Question ids to answers.

    Returns:
        Question ids to the hints their answers set off; quiet ones left out.
    """
    found: dict[str, tuple[Hint, ...]] = {}
    for question, answer in answers.items():
        items = [answer] if isinstance(answer, str) else list(answer)
        hints = tuple(hint for item in items for hint in review_answer(template, question, item))
        if hints:
            found[question] = hints
    return found

load_suite

load_suite(path: Path) -> Suite

Load a suite file and check all of it.

Parameters:

Name Type Description Default
path Path

The suite's TOML file.

required

Returns:

Type Description
Suite

The suite, with its prompt path resolved against the suite's directory.

Raises:

Type Description
SuiteError

If the file is not valid TOML or breaks the format, listing every problem at once.

OSError

If the file cannot be read.

Source code in src/validia/suites/suite.py
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
def load_suite(path: Path) -> Suite:
    """Load a suite file and check all of it.

    Args:
        path: The suite's TOML file.

    Returns:
        The suite, with its prompt path resolved against the suite's directory.

    Raises:
        SuiteError: If the file is not valid TOML or breaks the format, listing
            every problem at once.
        OSError: If the file cannot be read.
    """
    try:
        data = tomllib.loads(path.read_text(encoding="utf-8"))
    except tomllib.TOMLDecodeError as exc:
        msg = f"{path} is not valid TOML: {exc}"
        raise SuiteError(msg) from None

    reader = _Reader()
    reader.fields(data, "", ("description", "prompt", "tools", "grade", "cases"))
    description = reader.text(data, "description", "", required=False)
    prompt_name = reader.text(data, "prompt", "")
    prompt = path.parent / prompt_name
    if prompt_name and not prompt.is_file():
        reader.problems.append(f"prompt: no such file {prompt_name!r} next to the suite")
    grade = _grade(reader, data)
    tools_name = reader.text(data, "tools", "", required=grade is not None and grade.type == "tool")
    tools = _tools(reader, path.parent / tools_name, tools_name) if tools_name else ()
    cases = _cases(reader, data, grade, tools)

    if reader.problems or grade is None:
        _raise(path, reader.problems)
    return Suite(
        path=path,
        prompt=prompt,
        grade=grade,
        cases=cases,
        description=description,
        tools=tools,
        tools_file=path.parent / tools_name if tools_name else None,
    )