Skip to content
Verificate
Evidence & evaluation

Inspect the claim.
Understand its limits.

A result is only useful if you know what was measured, on what, and what it leaves unproven. This page sets that out for each result we publish. None of them replaces a test on your own questions and data.

Six different questions

“Does it work?” is several questions, not one.

These are often reported as a single accuracy figure. They are separate measurements, and a good result on one says little about the others.

  1. 01

    Does the citation resolve?

    The cited source exists and can be opened.

  2. 02

    Does the source support the claim?

    The evidence says what the answer says it does.

  3. 03

    Is the answer correct?

    The result matches a trusted reference.

  4. 04

    Does the decision repeat?

    The same inputs, records and rules select the same choice.

  5. 05

    How much gets answered?

    The share of useful questions that receive an answer, and how often the system declines.

  6. 06

    Is the outcome fair?

    Results meet a stated fairness test on a stated population.

Published results

What each result shows — and what it does not.

Every entry links to its source. Where a result comes from research rather than a customer deployment, it says so.

  • Helix

    Serving latency and valid structured output

    What was measured
    HELIX v1.7 compared with stock Ollama using the vLLM project’s GuideLLM harness on the same AMD EPYC 9254 machine, 32K context, at one and four concurrent requests.
    What it does not establish
    One model, one machine and one comparison system. It does not predict results on other hardware, models or serving stacks. The confidence features in v1.8 were added after these runs.
    Benchmarks
  • Helix

    A 120-billion-parameter model at interactive speed without a GPU

    What was measured
    GPT-OSS 120B throughput on three AMD Zen5 EPYC processors, CPU only, three-run average, run by AMD engineering in September 2026.
    What it does not establish
    Throughput only — it says nothing about answer quality. Results depend heavily on memory configuration, as the published table shows.
    Helix product page
  • Compile-Time Inference

    A continually updated decision policy can outperform a frozen one

    What was measured
    A study on a public procurement event log (BPI Challenge 2019), including a control that scrambles labels to test whether the gain is real.
    What it does not establish
    One dataset and one task. It does not show the same gain on your data, and it is a research result rather than a product guarantee.
    Research
  • Compile-Time Inference

    When a learned decision policy beats a fixed rule — and when it does not

    What was measured
    A randomised trial and sensitivity analysis identifying the conditions under which a learned policy outperforms a fixed business rule.
    What it does not establish
    The finding includes cases where a fixed rule is the better choice. It is specific to the studied setting.
    Research
  • Compile-Time Inference

    The production engine matches its reference implementation

    What was measured
    Zero mismatched selections in 10,000 comparisons between the production engine and the reference implementation.
    What it does not establish
    Shows the two implementations agree with each other. It does not show that either selects the correct decision.
    Research
  • Research paper

    Published method: decision policies

    What was measured
    Verificate EAV-DT, arXiv 2606.29280 (June 2026).
    What it does not establish
    A preprint. Read the paper’s own limitations section; results are specific to its workloads.
    Read on arXiv (opens a new tab)
  • Research paper

    Published method: detecting likely hallucination

    What was measured
    Hallucination prevention (CORTEX), arXiv 2602.17691 (January 2026).
    What it does not establish
    A preprint. Detection rates depend on the model and task, and no detector catches everything.
    Read on arXiv (opens a new tab)
  • Gate

    Gate caught planted defects that a model’s self-review missed

    What was measured
    Six deliberately planted defects: a frontier model reviewing its own work caught none; Gate caught all six.
    What it does not establish
    A small, adversarial test set built to probe a specific weakness. It is not a general defect-detection rate.
    Gate proof page

Customer evidence is product-specific. Helix is deployed at UNSW; that does not mean UNSW uses Compile-Time Inference or Gate.

Evaluate it yourself

The result that matters is the one on your workload.

A short, well-designed test tells you more than any published figure.

  1. 01

    Write the questions first

    Include routine, difficult and unanswerable questions, with reference answers, before you see any system’s output.

  2. 02

    Compare fairly

    Test against your current approach configured as well as you can configure it, on the same questions.

  3. 03

    Measure each question separately

    Report citation validity, claim support, correctness, repeatability, coverage and refusal on their own lines.

  4. 04

    Count the whole cost

    Include preparation, updates, storage, serving and review time — not only the price of a model response.