Open research

Evaluate AI against regional reality—not a translation of it.

We build native benchmarks for tasks where context changes the answer: freight operations, healthcare purchasing, financial risk, and customer support. Every case begins with authorized regional evidence and professional review.

Explore the research ↓
01Local evidenceDocuments, records, and rules from the market being evaluated
02Real tasksComplete decisions, not isolated knowledge questions
03Cost-aware errorsFailures weighted by their operational consequence
Preprint releaseSC—CO / v0.1.0
es-CO · FREIGHT OPERATIONS

Savia Carga · Colombia

An evaluation program for operational reasoning in road freight: local terminology, grounded calculations, exception handling, calibration, and safe abstention.

Source-groundedExpert reviewedes-CO nativeReproducible
Market
Colombia
Domain
Road freight operations
Unit
Operational decision
Evidence
Source required
Version 0.1.0 is published with its answer key, derivations, baseline result and sources. The questions have not yet been confirmed by a domain expert, and the page says so. Open the release ↗
What we measure
01

Source fidelity

Does the answer stay inside the evidence?

02

Operational correctness

Would the decision work in practice?

03

Calibrated abstention

Does the system know when not to answer?

Research agenda

Questions that matter in production.

01

Grounding

Can the system support every claim with authorized, current evidence?

02

Regional judgment

Does it recognize when a rule, term, or practice changes across countries?

03

Safe abstention

Does it stop and escalate when the evidence cannot support a responsible answer?

Methodology

A benchmark should explain why a system fails, not only how often.

Each release documents authorization, case origins, expert criteria, regional assumptions, and use limits. Results include error patterns and the evidence required to reproduce them.

  1. 01Regional task definition
  2. 02Authorized evidence set
  3. 03Expert scoring rubric
  4. 04Transparent error analysis