API reference¶
validia ¶
validia — a universal, async-first evaluation framework for LLM systems.
Covers the full evaluation spectrum under one set of primitives: prompt evaluation, tool selection, agent evaluation, and agent-type comparison.
The public API is re-exported from this module; everything not listed in
__all__ is internal and may change without a major version bump. It is the
same API whatever the front end: the validia command calls these functions,
and so can a REST service or a notebook -- see :mod:validia.suites.api.
CATEGORIES
module-attribute
¶
CATEGORIES = (
"wording",
"context",
"reasoning",
"output",
"tools",
"security",
"maintenance",
)
The core categories, in reading order.
CLOSING
module-attribute
¶
CLOSING = 'closing'
The id of the last question: how the model should answer.
PromptTemplate
dataclass
¶
A loaded, checked template.
Attributes:
| Name | Type | Description |
|---|---|---|
sections |
tuple[Section, ...]
|
The questions, in order. |
hints |
tuple[Hint, ...]
|
The core checks run on each answer. |
output |
dict[AnswerType, tuple[str, ...]]
|
For each answer type, the closing instructions to offer, first one the default. |
source |
str
|
Where the template came from, for the user to see. |
Source code in src/validia/prompts/template.py
140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 | |
sections_for ¶
sections_for(answer: AnswerType) -> tuple[Section, ...]
List the sections that apply to an answer type, in order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
answer
|
AnswerType
|
The suite's answer type. |
required |
Returns:
| Type | Description |
|---|---|
tuple[Section, ...]
|
The sections. |
Source code in src/validia/prompts/template.py
157 158 159 160 161 162 163 164 165 166 | |
hints_for ¶
hints_for(section: str, answer: str) -> list[Hint]
List the hints an answer sets off.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
section
|
str
|
The id of the section the answer is for. |
required |
answer
|
str
|
The answer. |
required |
Returns:
| Type | Description |
|---|---|
list[Hint]
|
The hints that fire, in template order. |
Source code in src/validia/prompts/template.py
168 169 170 171 172 173 174 175 176 177 178 | |
TemplateError ¶
Bases: ValueError
A prompt template is malformed; the message lists every problem in it.
Source code in src/validia/prompts/template.py
56 57 | |
Finding
dataclass
¶
One hit of one rule.
Attributes:
| Name | Type | Description |
|---|---|---|
rule |
str
|
The rule's id. |
severity |
Severity
|
How much it matters on the target model. |
line |
int
|
One-based line of the hit. |
column |
int
|
One-based column of the hit. |
excerpt |
str
|
The text that matched, shortened. |
fix |
str
|
What to do about it. |
Source code in src/validia/rules/model.py
265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 | |
Guidance
dataclass
¶
How to write for one target, in one category, with where the claims come from.
Attributes:
| Name | Type | Description |
|---|---|---|
category |
str
|
The category. |
target |
str
|
|
version |
str
|
The release its file comes from. |
summary |
str
|
What matters about the target, in a sentence. |
instructions |
tuple[str, ...]
|
What to do when writing a prompt for it. |
sources |
tuple[str, ...]
|
Where each claim comes from. |
Source code in src/validia/rules/model.py
286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 | |
Part
dataclass
¶
One released file of the core rules that was read.
Attributes:
| Name | Type | Description |
|---|---|---|
category |
str
|
The category. |
name |
str
|
The rule's folder name, or |
target |
str
|
|
version |
str
|
The release it changed in. |
released |
str
|
The day of that release. |
changes |
tuple[str, ...]
|
What changed. |
Source code in src/validia/rules/model.py
307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 | |
Release
dataclass
¶
One release of one category: every file that changed in it, and why.
Attributes:
| Name | Type | Description |
|---|---|---|
category |
str
|
The category. |
version |
str
|
The release, as in |
released |
str
|
The day of the release. |
changes |
tuple[tuple[str, tuple[str, ...]], ...]
|
Each change, with the files that carry it, as |
Source code in src/validia/rules/model.py
328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 | |
Rule
dataclass
¶
One check on prompt text.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Names the rule: its category and its folder, as in |
title |
str
|
What it finds, in a few words; the rule reference's heading. |
pattern |
str
|
What to look for. With |
fix |
str
|
What to do about a hit. |
severity |
Severity
|
How much a hit matters, on the models in |
scope |
tuple[Scope, ...]
|
The text it reads: the prompt, or tool descriptions. |
case_sensitive |
bool
|
Whether case matters, as it does for capitals. |
unless |
str
|
A hit is dropped when this matches in its sentence or the next. |
run |
int
|
Fire once on this many consecutive matching sentences or lines. |
unit |
Literal['sentence', 'line']
|
What |
at_least |
int
|
Fire once when the pattern matches this many times. |
missing |
bool
|
Fire when the pattern matches nowhere. |
models |
tuple[str, ...]
|
Model names (globs) the severity applies to; empty means all. |
otherwise |
Severity
|
The severity on every other model. |
confidence |
Literal['high', 'med', 'low']
|
How sure a hit is: |
judge |
str
|
The class set a judge would sort hits into, when one is needed. |
category |
str
|
The category the rule belongs to. |
fires |
tuple[str, ...]
|
Text the rule must fire on. |
quiet |
tuple[str, ...]
|
Text the rule must stay quiet on. |
origin |
tuple[str, ...]
|
Every file laid over it, in order, as anthropic/default 1.0.0 extend. |
Source code in src/validia/rules/model.py
121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 | |
severity_for ¶
severity_for(model: str | None) -> Severity
Say how much a hit matters on a model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
str | None
|
The model name, as in |
required |
Returns:
| Type | Description |
|---|---|
Severity
|
|
Source code in src/validia/rules/model.py
169 170 171 172 173 174 175 176 177 178 179 180 181 182 | |
describe_severity ¶
describe_severity(model: str | None) -> str
Say how much a hit matters: on one model, or across the models it names.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
str | None
|
The model name; |
required |
Returns:
| Type | Description |
|---|---|
str
|
|
Source code in src/validia/rules/model.py
184 185 186 187 188 189 190 191 192 193 194 195 | |
describe_matching ¶
describe_matching() -> str
Say in words how this rule turns matches into findings.
Returns:
| Type | Description |
|---|---|
str
|
A sentence, as in |
Source code in src/validia/rules/model.py
197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 | |
scan ¶
scan(text: str, model: str | None = None) -> list[Finding]
Find this rule's hits in a text.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The text to read. |
required |
model
|
str | None
|
The target model, for the severity. |
None
|
Returns:
| Type | Description |
|---|---|
list[Finding]
|
The findings, in order. |
Source code in src/validia/rules/model.py
217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 | |
RuleCheck
dataclass
¶
The outcome of proving one rule, or one example, against its cases.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The rule's id, or the example's name. |
failures |
tuple[str, ...]
|
What went wrong; empty when it passed. |
cases |
int
|
How many cases were run. |
category |
str
|
The category it belongs to. |
Source code in src/validia/rules/model.py
440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 | |
RuleError ¶
Bases: ValueError
Rules are malformed; problems lists every reason.
Attributes:
| Name | Type | Description |
|---|---|---|
source |
The file or folder the problems are in. |
|
problems |
One line per problem. |
Source code in src/validia/rules/model.py
104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 | |
__init__ ¶
__init__(source: str, problems: Sequence[str]) -> None
Keep the problems, and say them all in the message.
Source code in src/validia/rules/model.py
112 113 114 115 116 117 118 | |
RulePack
dataclass
¶
A loaded, checked set of rules for one target.
Attributes:
| Name | Type | Description |
|---|---|---|
rules |
tuple[Rule, ...]
|
The rules, category by category. |
examples |
tuple[Example, ...]
|
Whole-prompt examples that still hold for these rules. |
parts |
tuple[Part, ...]
|
Every core file read. |
guidance |
tuple[Guidance, ...]
|
What the vendor and model files say about writing for the target. |
releases |
tuple[tuple[str, str], ...]
|
Each category, and the release it was read at. |
target |
Target | None
|
The model the rules were resolved for, if any. |
project |
str
|
The project folder laid over them, if any. |
Source code in src/validia/rules/model.py
370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 | |
version
property
¶
version: str
Name these rules as a report records them, so a result can be reproduced.
Returns:
| Type | Description |
|---|---|
str
|
The common release, then any category on another one, then the target and |
str
|
the project folder: |
only ¶
only(categories: Sequence[str]) -> RulePack
Keep only some categories.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
categories
|
Sequence[str]
|
The categories to keep. |
required |
Returns:
| Type | Description |
|---|---|
RulePack
|
The narrower pack. |
Source code in src/validia/rules/model.py
420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 | |
Written
dataclass
¶
What a write left in the project, and the proof that it holds.
Attributes:
| Name | Type | Description |
|---|---|---|
paths |
tuple[Path, ...]
|
The files written, the rule's first, then its cases if any. |
pack |
RulePack
|
The project's rules for the target, as read after the write. |
checks |
tuple[RuleCheck, ...]
|
The rule's own check, and every example that names it. |
Source code in src/validia/rules/local.py
75 76 77 78 79 80 81 82 83 84 85 86 87 | |
RunSummary
dataclass
¶
A run, summed up.
Attributes:
| Name | Type | Description |
|---|---|---|
trials |
int
|
Trials run. |
passed |
int
|
Trials that passed. |
graded |
int
|
Trials that got a reply to grade. |
errors |
Mapping[str, int]
|
Failure classes of the trials whose call failed, and how many of each. |
interval |
tuple[float, float]
|
The 95% Wilson interval of the pass rate, as fractions. |
cases |
tuple[CaseResult, ...]
|
Every case, in suite order. |
groups |
Mapping[str, tuple[int, int]]
|
Each group's |
input_tokens |
int
|
Tokens read, across the run. |
output_tokens |
int
|
Tokens written, across the run. |
cache_read_tokens |
int
|
Input tokens served from cache, across the run. |
latency_p50_ms |
float
|
The median latency of the calls that answered. |
latency_p95_ms |
float
|
Their 95th percentile. |
served_models |
tuple[str, ...]
|
Every model the provider says answered. |
Source code in src/validia/runs/runner.py
407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 | |
to_dict ¶
to_dict() -> dict[str, Any]
Render the summary as plain data, for summary.json.
Source code in src/validia/runs/runner.py
446 447 448 449 450 451 | |
Trial
dataclass
¶
One case, sent once.
Attributes:
| Name | Type | Description |
|---|---|---|
case |
Case
|
The case. |
rep |
int
|
Which repetition this is, from 1. |
Source code in src/validia/runs/runner.py
215 216 217 218 219 220 221 222 223 224 225 | |
TrialResult
dataclass
¶
What one trial came to.
Attributes:
| Name | Type | Description |
|---|---|---|
case |
str
|
The case's id. |
rep |
int
|
Which repetition, from 1. |
tags |
tuple[str, ...]
|
The case's tags; the first is its group. |
passed |
bool | None
|
Whether the reply was right; |
reason |
str
|
Why it failed: the grader's reason, or the call's error. Empty on a pass. |
reply |
str
|
The reply's text, for the record. |
tool_calls |
tuple[str, ...]
|
The tools it called, as |
error |
str | None
|
The failure class of a call that failed, as franca names it. |
input_tokens |
int
|
Tokens the call read. |
output_tokens |
int
|
Tokens it wrote. |
cache_read_tokens |
int
|
Input tokens served from the provider's cache. |
latency_ms |
float
|
How long the last attempt took on the wire. |
attempts |
int
|
How many calls the trial took, retries included. |
served_model |
str | None
|
The model the provider says answered. |
stop_reason |
str | None
|
Why it stopped, as the provider says. |
Source code in src/validia/runs/runner.py
241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 | |
to_json ¶
to_json() -> str
Render the result as one line of JSON, for trials.jsonl.
Source code in src/validia/runs/runner.py
279 280 281 | |
BuildError ¶
Bases: ValueError
The answers cannot make a prompt; problems lists every reason.
Attributes:
| Name | Type | Description |
|---|---|---|
problems |
One line per problem, as in |
Source code in src/validia/prompts/building.py
40 41 42 43 44 45 46 47 48 49 50 | |
__init__ ¶
__init__(problems: Sequence[str]) -> None
Keep the problems, and say them all in the message.
Source code in src/validia/prompts/building.py
47 48 49 50 | |
CaseSpec
dataclass
¶
A case to add: one input, and what a right answer looks like.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Unique within the suite; letters, digits, |
input |
str
|
The message the model gets. |
expected |
Any
|
In the suite's answer type's shape: a label; a dict of JSON
fields; a dict of text checks; or |
tags |
tuple[str, ...]
|
Ordered labels; the first is the group results are reported by. |
Source code in src/validia/suites/api.py
128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 | |
from_dict
classmethod
¶
from_dict(
data: Mapping[str, Any], where: str = ""
) -> CaseSpec
Build a case from a JSON object, as a request body carries it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Mapping[str, Any]
|
The object. |
required |
where
|
str
|
A prefix for problem locations, such as |
''
|
Returns:
| Type | Description |
|---|---|
CaseSpec
|
The case. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If a field is unknown, missing or of the wrong type. |
Source code in src/validia/suites/api.py
145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 | |
problems ¶
problems(
taken: Collection[str] = (), where: str = ""
) -> list[str]
Check what can be checked without the suite file.
The expected answer's shape is checked against the suite itself when the case is written, by the same loader that reads every suite.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
taken
|
Collection[str]
|
Ids already in use. |
()
|
where
|
str
|
A prefix for problem locations. |
''
|
Returns:
| Type | Description |
|---|---|
list[str]
|
Every problem found. |
Source code in src/validia/suites/api.py
173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 | |
Option
dataclass
¶
One answer a question offers.
Attributes:
| Name | Type | Description |
|---|---|---|
value |
str
|
The answer itself, as it is written into the prompt. |
label |
str
|
A short description, when the value does not speak for itself. |
Source code in src/validia/prompts/building.py
53 54 55 56 57 58 59 60 61 62 63 | |
Question
dataclass
¶
One thing to ask, in a form any front end can show.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Where the answer goes: a section id, or :data: |
text |
str
|
The question. |
kind |
Literal['text', 'many', 'choice']
|
|
required |
bool
|
Whether an empty answer is refused, or skips the part. |
example |
str
|
An answer that shows the expected shape. |
options |
tuple[Option, ...]
|
For |
custom |
bool
|
For |
Source code in src/validia/prompts/building.py
66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 | |
SpecError ¶
Bases: ValueError
A request cannot be carried out; problems lists every reason.
Attributes:
| Name | Type | Description |
|---|---|---|
problems |
One line per problem, as in |
Source code in src/validia/suites/api.py
86 87 88 89 90 91 92 93 94 95 96 | |
__init__ ¶
__init__(problems: Sequence[str]) -> None
Keep the problems, and say them all in the message.
Source code in src/validia/suites/api.py
93 94 95 96 | |
SuiteSpec
dataclass
¶
A suite to create: where it goes, how it is graded, its prompt and its cases.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The suite's folder name. |
answer |
AnswerType
|
What the prompt answers with: |
cases |
tuple[CaseSpec, ...]
|
At least one case. |
folder |
str
|
The folder suites live in, relative to the project. |
labels |
tuple[str, ...]
|
For |
required |
tuple[str, ...]
|
For |
prompt |
str
|
The prompt text, written to |
prompt_file |
str
|
An existing prompt, relative to the project, that the suite points at instead of copying. |
tools |
tuple[Tool, ...]
|
For |
tools_file |
str
|
An existing tools file, relative to the project. |
description |
str
|
What the suite measures. |
Source code in src/validia/suites/api.py
215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 | |
from_dict
classmethod
¶
from_dict(data: Mapping[str, Any]) -> SuiteSpec
Build a spec from a JSON object, as a request body carries it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Mapping[str, Any]
|
The object. |
required |
Returns:
| Type | Description |
|---|---|
SuiteSpec
|
The spec, its fields typed but not yet checked against each other; |
SuiteSpec
|
func: |
Raises:
| Type | Description |
|---|---|
SpecError
|
If a field is unknown, missing or of the wrong type, listing every problem at once. |
Source code in src/validia/suites/api.py
251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 | |
problems ¶
problems() -> list[str]
Check the spec's parts against each other, without touching a disk.
Returns:
| Type | Description |
|---|---|
list[str]
|
Every problem found. |
Source code in src/validia/suites/api.py
312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 | |
Reply
dataclass
¶
Everything the model answered with, for one case.
Attributes:
| Name | Type | Description |
|---|---|---|
text |
str
|
The text of the answer. |
tool_calls |
tuple[ToolCall, ...]
|
Tools it asked to call, in the order it asked. |
Source code in src/validia/suites/expect.py
50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 | |
ToolCall
dataclass
¶
One tool the model asked to call.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The tool's name. |
args |
dict[str, Any]
|
Its arguments, decoded. |
Source code in src/validia/suites/expect.py
37 38 39 40 41 42 43 44 45 46 47 | |
Verdict
dataclass
¶
The outcome of checking one reply.
Attributes:
| Name | Type | Description |
|---|---|---|
passed |
bool
|
Whether the reply is right. |
reason |
str
|
Why it is not, when it is not; empty on a pass. |
Source code in src/validia/suites/expect.py
68 69 70 71 72 73 74 75 76 77 78 | |
Case
dataclass
¶
One input to run the prompt on, and what a right answer looks like.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Unique within the suite; results and transcripts are keyed by it. |
input |
str
|
The user message sent to the model. |
expected |
Expected
|
What the reply is checked against. |
tags |
tuple[str, ...]
|
Ordered labels; the first is the group results are reported by. |
Source code in src/validia/suites/suite.py
130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 | |
check ¶
check(reply: Reply) -> Verdict
Grade one reply to this case.
An empty reply -- no text and no tool call -- always fails, whatever the answer type: no answer is not a wrong answer, and it must never score like a considered one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
reply
|
Reply
|
The model's answer. |
required |
Returns:
| Type | Description |
|---|---|
Verdict
|
The verdict. |
Source code in src/validia/suites/suite.py
146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 | |
Grade
dataclass
¶
How each reply is scored.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
AnswerType
|
The answer type, which decides the shape of every |
labels |
tuple[str, ...]
|
For |
required |
tuple[str, ...]
|
For |
Source code in src/validia/suites/suite.py
98 99 100 101 102 103 104 105 106 107 108 109 110 | |
Suite
dataclass
¶
A loaded, validated suite.
Attributes:
| Name | Type | Description |
|---|---|---|
path |
Path
|
The suite file. |
prompt |
Path
|
The prompt file, resolved against the suite's directory. |
grade |
Grade
|
How replies are scored. |
cases |
tuple[Case, ...]
|
The cases, in file order. |
description |
str
|
What the suite measures. |
tools |
tuple[Tool, ...]
|
The tools the model is offered, if any. |
tools_file |
Path | None
|
The file they were read from. |
Source code in src/validia/suites/suite.py
164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 | |
summary ¶
summary() -> str
Describe the grader in one line, as run --dry-run prints it.
Returns:
| Type | Description |
|---|---|
str
|
The answer type, and what it checks against. |
Source code in src/validia/suites/suite.py
186 187 188 189 190 191 192 193 194 195 196 197 198 199 | |
groups ¶
groups() -> dict[str, int]
Count cases by their first tag, the group results are reported by.
Returns:
| Type | Description |
|---|---|
dict[str, int]
|
Case counts, in order of first appearance. |
Source code in src/validia/suites/suite.py
201 202 203 204 205 206 207 | |
SuiteError ¶
Bases: ValueError
A suite file is malformed; the message lists every problem in it.
Attributes:
| Name | Type | Description |
|---|---|---|
problems |
One line per problem, for a caller that reports them as data. |
Source code in src/validia/suites/suite.py
85 86 87 88 89 90 91 92 93 94 95 | |
__init__ ¶
__init__(
message: str, problems: Sequence[str] = ()
) -> None
Keep the problems beside the rendered message.
Source code in src/validia/suites/suite.py
92 93 94 95 | |
Tool
dataclass
¶
A tool the model is offered, as franca's ToolDef takes it.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The tool's name, as the model will spell it in a call. |
parameters |
dict[str, Any]
|
The JSON Schema of its arguments. |
description |
str
|
What it does, in the model's own context window. |
strict |
bool | None
|
Whether the provider must enforce the schema exactly. |
Source code in src/validia/suites/suite.py
113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | |
default_template ¶
default_template() -> PromptTemplate
Load the built-in template.
Returns:
| Type | Description |
|---|---|
PromptTemplate
|
The template. |
Source code in src/validia/prompts/template.py
378 379 380 381 382 383 384 | |
load_template ¶
load_template(path: Path) -> PromptTemplate
Load a project's template.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The |
required |
Returns:
| Type | Description |
|---|---|
PromptTemplate
|
The template. |
Raises:
| Type | Description |
|---|---|
TemplateError
|
If it is malformed, listing every problem. |
OSError
|
If it cannot be read. |
Source code in src/validia/prompts/template.py
387 388 389 390 391 392 393 394 395 396 397 398 399 400 | |
core_categories ¶
core_categories(
folder: Traversable | Path | None = None,
) -> tuple[str, ...]
List the core categories, in reading order, then any others by name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[str, ...]
|
The category names. |
Source code in src/validia/rules/catalog.py
63 64 65 66 67 68 69 70 71 72 73 74 | |
core_releases ¶
core_releases(
category: str, folder: Traversable | Path | None = None
) -> tuple[Release, ...]
List a category's releases, oldest first, with what changed in each.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
category
|
str
|
The category. |
required |
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[Release, ...]
|
The releases. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If a file's header is malformed, or a release's files disagree on the day it was released. |
Source code in src/validia/rules/catalog.py
160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 | |
core_rules ¶
core_rules(
category: str, folder: Traversable | Path | None = None
) -> tuple[str, ...]
List a category's rule folders, by name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
category
|
str
|
The category. |
required |
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[str, ...]
|
Rule names, such as |
Source code in src/validia/rules/catalog.py
77 78 79 80 81 82 83 84 85 86 87 | |
core_targets ¶
core_targets(
category: str, folder: Traversable | Path | None = None
) -> tuple[str, ...]
List every target any file of a category is for: the default, then by vendor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
category
|
str
|
The category. |
required |
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[str, ...]
|
Targets such as |
Source code in src/validia/rules/catalog.py
97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 | |
core_versions ¶
core_versions(
category: str,
name: str,
target: str = "default",
folder: Traversable | Path | None = None,
) -> tuple[str, ...]
List the releases one rule, or a category's examples or guidance, changed in.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
category
|
str
|
The category. |
required |
name
|
str
|
The rule's folder name, or |
required |
target
|
str
|
The target folder. |
'default'
|
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[str, ...]
|
Versions, oldest first; empty when there is no such folder. |
Source code in src/validia/rules/catalog.py
128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | |
default_rules ¶
default_rules(
target: Target | None = None,
*,
pins: Mapping[str, str] | None = None,
categories: Sequence[str] | None = None,
folder: Traversable | Path | None = None,
) -> RulePack
Load the core rules for a target: every rule, read along its chain.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
target
|
Target | None
|
The model, as |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the release to read; |
None
|
categories
|
Sequence[str] | None
|
The categories to keep; every one when omitted. |
None
|
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
RulePack
|
The combined pack. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If a pin matches no release, or a file is malformed. |
Source code in src/validia/rules/layering.py
390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 | |
lint_text ¶
lint_text(
text: str,
pack: RulePack,
*,
model: str | None = None,
scope: Scope = "prompt",
) -> list[Finding]
Run every rule for a kind of text over it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The prompt, or one tool's description. |
required |
pack
|
RulePack
|
The rules. |
required |
model
|
str | None
|
The model name severities are read for; the pack's own when omitted. |
None
|
scope
|
Scope
|
What kind of text it is. |
'prompt'
|
Returns:
| Type | Description |
|---|---|
list[Finding]
|
Every finding, ordered by position. |
Source code in src/validia/rules/matching.py
8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 | |
load_rules ¶
load_rules(
path: Path,
*,
target: Target | None = None,
pins: Mapping[str, str] | None = None,
categories: Sequence[str] | None = None,
folder: Traversable | Path | None = None,
) -> RulePack
Load a project's rules/ folder, laid over the core rules for a target.
The folder has the core's layout and is read by the same code, after the core: each rule's chain runs through the core's default, vendor and model files, then the project's. Every file in the folder is checked, whichever categories are kept, and every problem is reported at once.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The project's |
required |
target
|
Target | None
|
The model, as |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the core release to read; |
None
|
categories
|
Sequence[str] | None
|
The categories to keep, the project's own included. |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
RulePack
|
The combined pack. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If either side is malformed, listing every problem in the folder. |
OSError
|
If a file cannot be read. |
Source code in src/validia/rules/layering.py
415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 | |
verify_core ¶
verify_core(
*,
pins: Mapping[str, str] | None = None,
categories: Sequence[str] | None = None,
targets: Sequence[str] | None = None,
folder: Traversable | Path | None = None,
) -> list[tuple[str, str, RulePack, list[RuleCheck]]]
Prove every core target in its own chain: each as it would be linted.
A vendor's files are proved over the default; a model's over both. Each examples file runs as its own model sees the rules, so a model's severities are tested where they are written.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pins
|
Mapping[str, str] | None
|
Category names to the release to read; |
None
|
categories
|
Sequence[str] | None
|
The categories to prove; every core category when omitted. |
None
|
targets
|
Sequence[str] | None
|
The targets to prove, as |
None
|
folder
|
Traversable | Path | None
|
Where the rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
list[tuple[str, str, RulePack, list[RuleCheck]]]
|
|
Source code in src/validia/rules/layering.py
596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 | |
verify_project ¶
verify_project(
path: Path,
*,
pins: Mapping[str, str] | None = None,
folder: Traversable | Path | None = None,
) -> list[tuple[str, RulePack, list[RuleCheck]]]
Prove a project's rules/ folder for every target it has files for.
The default, then each vendor's and each model's, each read over the core, so a project's model-specific cases and examples are proved on their own model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The project's |
required |
pins
|
Mapping[str, str] | None
|
Category names to the core release to read; |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
list[tuple[str, RulePack, list[RuleCheck]]]
|
|
Raises:
| Type | Description |
|---|---|
RuleError
|
If the folder or the core is malformed. |
Source code in src/validia/rules/layering.py
561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 | |
verify_rules ¶
verify_rules(pack: RulePack) -> list[RuleCheck]
Prove every rule against its own cases, then every example.
An example is linted with its own category's rules, as its own model sees them, so a layer's examples prove the severities that layer sets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pack
|
RulePack
|
The rules. |
required |
Returns:
| Type | Description |
|---|---|
list[RuleCheck]
|
One check per rule, then one per example. |
Source code in src/validia/rules/matching.py
33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 | |
disable_rule ¶
disable_rule(
project: Path,
rule_id: str,
*,
target: str = "default",
version: str | None = None,
pins: Mapping[str, str] | None = None,
folder: Traversable | Path | None = None,
) -> Written
Turn a rule off for a target and everything read after it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
project
|
Path
|
The project's |
required |
rule_id
|
str
|
The rule, as |
required |
target
|
str
|
|
'default'
|
version
|
str | None
|
The file's version; the first free one, |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the core release to read. |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
Written
|
What was written; it proves nothing of the rule, which is gone. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If the rule is unknown, or the file exists; nothing is left written. |
Source code in src/validia/rules/local.py
387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 | |
extend_rule ¶
extend_rule(
project: Path,
rule_id: str,
*,
target: str = "default",
fields: Mapping[str, Any] | None = None,
fires: Sequence[str] = (),
quiet: Sequence[str] = (),
version: str | None = None,
pins: Mapping[str, str] | None = None,
folder: Traversable | Path | None = None,
) -> Written
Extend a rule for a target: change some fields, add cases, keep the rest.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
project
|
Path
|
The project's |
required |
rule_id
|
str
|
The rule, as |
required |
target
|
str
|
|
'default'
|
fields
|
Mapping[str, Any] | None
|
The fields to change, as a rule file names them ( |
None
|
fires
|
Sequence[str]
|
More text the rule must fire on. |
()
|
quiet
|
Sequence[str]
|
More text the rule must stay quiet on. |
()
|
version
|
str | None
|
The file's version; the first free one, |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the core release to read. |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
Written
|
What was written, and its proof. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If there is nothing to change, the rule is unknown, the file exists, or the result does not load or prove; nothing is left written. |
Source code in src/validia/rules/local.py
267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 | |
fall_back ¶
fall_back(
project: Path,
rule_id: str,
release: str,
*,
version: str | None = None,
pins: Mapping[str, str] | None = None,
folder: Traversable | Path | None = None,
) -> Written
Take one core rule, along its whole chain, as an older release had it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
project
|
Path
|
The project's |
required |
rule_id
|
str
|
The rule, as |
required |
release
|
str
|
The core release, as |
required |
version
|
str | None
|
The file's version; the first free one, |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the core release to read for everything else. |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
Written
|
What was written, and its proof. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If the release or rule is unknown, or the file exists; nothing is left written. |
Source code in src/validia/rules/local.py
417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 | |
new_rule ¶
new_rule(
project: Path,
rule_id: str,
*,
fields: Mapping[str, Any],
fires: Sequence[str],
quiet: Sequence[str],
target: str = "default",
version: str | None = None,
pins: Mapping[str, str] | None = None,
folder: Traversable | Path | None = None,
) -> Written
Add a rule of the project's own, in a core category or one of its own.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
project
|
Path
|
The project's |
required |
rule_id
|
str
|
The new rule, as |
required |
fields
|
Mapping[str, Any]
|
Its fields; |
required |
fires
|
Sequence[str]
|
Text it must fire on; at least one. |
required |
quiet
|
Sequence[str]
|
Text it must stay quiet on; at least one. |
required |
target
|
str
|
|
'default'
|
version
|
str | None
|
The file's version; the first free one, |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the core release to read. |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
Written
|
What was written, and its proof. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If the rule exists already, a field is missing or wrong, or its cases fail; nothing is left written. |
Source code in src/validia/rules/local.py
447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 | |
replace_rule ¶
replace_rule(
project: Path,
rule_id: str,
*,
target: str = "default",
version: str | None = None,
pins: Mapping[str, str] | None = None,
folder: Traversable | Path | None = None,
) -> Written
Copy a rule, as the target reads it now, into a file of its own to edit.
The copy is a full definition, so it replaces the rule from the target up: every field that differs from the defaults, and every case the rule has.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
project
|
Path
|
The project's |
required |
rule_id
|
str
|
The rule, as |
required |
target
|
str
|
|
'default'
|
version
|
str | None
|
The file's version; the first free one, |
None
|
pins
|
Mapping[str, str] | None
|
Category names to the core release to read. |
None
|
folder
|
Traversable | Path | None
|
Where the core rules live; the package's own when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
Written
|
What was written, and its proof. |
Raises:
| Type | Description |
|---|---|
RuleError
|
If the rule is unknown, the file exists, or the copy does not load or prove; nothing is left written. |
Source code in src/validia/rules/local.py
325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 | |
rule_target ¶
rule_target(target: str) -> Target | None
Turn a target folder into the model it is read for.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
target
|
str
|
|
required |
Returns:
| Type | Description |
|---|---|
Target | None
|
|
Target | None
|
default reads as the model |
Raises:
| Type | Description |
|---|---|
RuleError
|
If the target is not one of those shapes. |
Source code in src/validia/rules/local.py
90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 | |
build_model ¶
build_model(
provider: str,
model: str,
*,
keys: KeyProvider,
transport: Transport,
clock: Clock,
settings: ProviderSettings | None = None,
) -> ChatModel
Assemble franca's chat model for one provider and model id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
The provider's slug, as in |
required |
model
|
str
|
The model id to send, as in |
required |
keys
|
KeyProvider
|
Where the API key is read from, at call time. |
required |
transport
|
Transport
|
The HTTP implementation; a scripted one in tests. |
required |
clock
|
Clock
|
Source of time for latency and retry delays. |
required |
settings
|
ProviderSettings | None
|
|
None
|
Returns:
| Type | Description |
|---|---|
ChatModel
|
The model, ready for |
Raises:
| Type | Description |
|---|---|
AccessError
|
If franca has no chat endpoint for the provider. |
Source code in src/validia/runs/runner.py
92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | |
run_trial
async
¶
run_trial(
client: ChatClient,
suite: Suite,
prompt: str,
trial: Trial,
*,
clock: Clock,
retries: int = 2,
) -> TrialResult
Send one trial, retrying what is worth retrying, and grade the reply.
A failure franca marks retryable -- a rate limit, an overloaded server, a dropped connection -- is retried after the delay the provider asked for, or after an exponential backoff when it asked for none. Anything else, and a retryable failure that outlives its retries, ends the trial as an error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
ChatClient
|
The model, or a pipeline of middleware around it. |
required |
suite
|
Suite
|
The suite, for its tools. |
required |
prompt
|
str
|
The prompt under test. |
required |
trial
|
Trial
|
The trial. |
required |
clock
|
Clock
|
Where retry delays are slept; injected, so tests never wait. |
required |
retries
|
int
|
Extra attempts a retryable failure gets. |
2
|
Returns:
| Type | Description |
|---|---|
TrialResult
|
The trial's result. A failed call is a result, not an exception. |
Source code in src/validia/runs/runner.py
284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 | |
summarize ¶
summarize(results: Sequence[TrialResult]) -> RunSummary
Sum up a run's trial results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
Sequence[TrialResult]
|
Every trial's result, in trial order. |
required |
Returns:
| Type | Description |
|---|---|
RunSummary
|
The summary: overall, by case, by group, and the cost. |
Source code in src/validia/runs/runner.py
454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 | |
trials ¶
trials(suite: Suite, reps: int) -> list[Trial]
Every trial a run makes: each case, reps times, case by case.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
suite
|
Suite
|
The suite. |
required |
reps
|
int
|
Repetitions of every case; at least 1. |
required |
Returns:
| Type | Description |
|---|---|
list[Trial]
|
The trials, in the order they are reported. |
Source code in src/validia/runs/runner.py
228 229 230 231 232 233 234 235 236 237 238 | |
unsupported ¶
unsupported(suite: Suite) -> str | None
Say why a suite cannot be run through franca yet, if it cannot.
franca's adapters send text turns only, for now: a package's tools never reach the wire, and a tool call in the reply is not read back. A tool suite run anyway would grade every case as "called no tool" -- a wrong answer the model never gave -- so it is refused instead, until franca carries tools and validia's floor on it moves up.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
suite
|
Suite
|
The suite. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
The reason, or |
Source code in src/validia/runs/runner.py
189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 | |
wilson ¶
wilson(
passed: int, total: int, z: float = 1.96
) -> tuple[float, float]
The Wilson score interval for a pass rate: honest at 0%, 100% and small n.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
passed
|
int
|
Trials that passed. |
required |
total
|
int
|
Trials graded. |
required |
z
|
float
|
The normal quantile; 1.96 gives a 95% interval. |
1.96
|
Returns:
| Type | Description |
|---|---|
tuple[float, float]
|
|
Source code in src/validia/runs/runner.py
358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 | |
add_case ¶
add_case(suite_path: Path, case: CaseSpec) -> Suite
Append one case to a suite, keeping it only if the suite still loads.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
suite_path
|
Path
|
The suite file. |
required |
case
|
CaseSpec
|
The case. |
required |
Returns:
| Type | Description |
|---|---|
Suite
|
The suite as it now stands. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If the case does not fit the suite, listing every problem. |
Source code in src/validia/suites/api.py
483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 | |
check_reply ¶
check_reply(
suite: Suite, case_id: str, reply: Reply
) -> Verdict
Grade one reply against one case, without calling a model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
suite
|
Suite
|
The suite. |
required |
case_id
|
str
|
The case's id. |
required |
reply
|
Reply
|
The reply to grade. |
required |
Returns:
| Type | Description |
|---|---|
Verdict
|
The verdict. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If there is no such case, suggesting the closest one. |
Source code in src/validia/suites/api.py
562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 | |
closing_options ¶
closing_options(
template: PromptTemplate,
grade: Grade,
tools: Sequence[Tool],
) -> list[str]
List the closing instructions on offer, filled in for this suite.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
template
|
PromptTemplate
|
Where the instructions come from. |
required |
grade
|
Grade
|
The suite's grading: its answer type, labels and keys. |
required |
tools
|
Sequence[Tool]
|
The suite's tools. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
The instructions, the template's default first. |
Source code in src/validia/prompts/building.py
90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 | |
create_suite ¶
create_suite(spec: SuiteSpec, root: Path) -> Suite
Create a suite: its file, its prompt and tools when given as text, all at once.
Nothing is written unless the whole suite loads; a failure leaves no trace.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
spec
|
SuiteSpec
|
What to create. |
required |
root
|
Path
|
The project the spec's folders and files are relative to. |
required |
Returns:
| Type | Description |
|---|---|
Suite
|
The suite, as loaded from what was written. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If the spec, a referenced file, or the written suite has a problem, listing every one. |
Source code in src/validia/suites/api.py
401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 | |
describe_suite ¶
describe_suite(
suite: Suite, root: Path | None = None
) -> dict[str, Any]
Describe a suite as JSON-ready data: what a front end needs to show it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
suite
|
Suite
|
The suite. |
required |
root
|
Path | None
|
Paths are shown relative to this, when they are inside it. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
The answer type, labels, required keys, tools with their parameters, and |
dict[str, Any]
|
every case's id, tags and input. |
Source code in src/validia/suites/api.py
579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 | |
find_case ¶
find_case(suite: Suite, case_id: str) -> Case
Look a case up by id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
suite
|
Suite
|
The suite. |
required |
case_id
|
str
|
The case's id. |
required |
Returns:
| Type | Description |
|---|---|
Case
|
The case. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If there is no such case, suggesting the closest one. |
Source code in src/validia/suites/api.py
541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 | |
locate_suite ¶
locate_suite(
root: Path, name: str, folder: str = "evals"
) -> Path
Find a suite file by name, refusing any name that would leave the project.
A front end that takes a suite's name from a URL or a form should come through
here rather than joining paths itself: .. in a name is a request for a file
the API was never meant to touch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
root
|
Path
|
The project. |
required |
name
|
str
|
The suite's name. |
required |
folder
|
str
|
The folder suites live in, relative to the project. |
'evals'
|
Returns:
| Type | Description |
|---|---|
Path
|
The suite file's path. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If the name or folder is not a plain path inside the project, or there is no such suite. |
Source code in src/validia/suites/api.py
365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 | |
parse_reply ¶
parse_reply(data: Mapping[str, Any]) -> Reply
Build a reply from a JSON object: {"text": ..., "tool_calls": [...]}.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Mapping[str, Any]
|
The object; each tool call is |
required |
Returns:
| Type | Description |
|---|---|
Reply
|
The reply. |
Raises:
| Type | Description |
|---|---|
SpecError
|
If a field is unknown or of the wrong type. |
Source code in src/validia/suites/api.py
506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 | |
prompt_questions ¶
prompt_questions(
template: PromptTemplate,
grade: Grade,
tools: Sequence[Tool] = (),
) -> tuple[Question, ...]
List the questions that build a prompt for this suite, in order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
template
|
PromptTemplate
|
The sections to ask about. |
required |
grade
|
Grade
|
The suite's grading, which picks the sections and the closing. |
required |
tools
|
Sequence[Tool]
|
The suite's tools, listed in a tool prompt's closing. |
()
|
Returns:
| Type | Description |
|---|---|
tuple[Question, ...]
|
One question per section that applies, then :data: |
Source code in src/validia/prompts/building.py
107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | |
render_prompt ¶
render_prompt(
template: PromptTemplate,
grade: Grade,
answers: Mapping[str, Answer],
tools: Sequence[Tool] = (),
) -> str
Turn a complete set of answers into the prompt.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
template
|
PromptTemplate
|
The sections the answers fill. |
required |
grade
|
Grade
|
The suite's grading, which picks the sections that apply. |
required |
answers
|
Mapping[str, Answer]
|
Question ids, as :func: |
required |
tools
|
Sequence[Tool]
|
The suite's tools. |
()
|
Returns:
| Type | Description |
|---|---|
str
|
The prompt, sections separated by blank lines, ending in a newline. |
Raises:
| Type | Description |
|---|---|
BuildError
|
If an answer is missing, misshapen, or for no question, listing every problem at once. |
Source code in src/validia/prompts/building.py
201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 | |
review_answer ¶
review_answer(
template: PromptTemplate, question: str, answer: str
) -> tuple[Hint, ...]
Run the template's core checks on one answer.
A hint is advice, never a refusal: the front end shows it and the person decides whether to rewrite.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
template
|
PromptTemplate
|
Where the checks come from. |
required |
question
|
str
|
The id of the question the answer is for. |
required |
answer
|
str
|
The answer. |
required |
Returns:
| Type | Description |
|---|---|
tuple[Hint, ...]
|
The hints it sets off, in template order. |
Source code in src/validia/prompts/building.py
144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | |
review_answers ¶
review_answers(
template: PromptTemplate, answers: Mapping[str, Answer]
) -> dict[str, tuple[Hint, ...]]
Run the core checks on a whole set of answers, as a form submits them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
template
|
PromptTemplate
|
Where the checks come from. |
required |
answers
|
Mapping[str, Answer]
|
Question ids to answers. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, tuple[Hint, ...]]
|
Question ids to the hints their answers set off; quiet ones left out. |
Source code in src/validia/prompts/building.py
161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 | |
load_suite ¶
load_suite(path: Path) -> Suite
Load a suite file and check all of it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The suite's TOML file. |
required |
Returns:
| Type | Description |
|---|---|
Suite
|
The suite, with its prompt path resolved against the suite's directory. |
Raises:
| Type | Description |
|---|---|
SuiteError
|
If the file is not valid TOML or breaks the format, listing every problem at once. |
OSError
|
If the file cannot be read. |
Source code in src/validia/suites/suite.py
479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 | |