Inspect the claim.
Understand its limits.
A result is only useful if you know what was measured, on what, and what it leaves unproven. This page sets that out for each result we publish. None of them replaces a test on your own questions and data.
“Does it work?” is several questions, not one.
These are often reported as a single accuracy figure. They are separate measurements, and a good result on one says little about the others.
- 01
Does the citation resolve?
The cited source exists and can be opened.
- 02
Does the source support the claim?
The evidence says what the answer says it does.
- 03
Is the answer correct?
The result matches a trusted reference.
- 04
Does the decision repeat?
The same inputs, records and rules select the same choice.
- 05
How much gets answered?
The share of useful questions that receive an answer, and how often the system declines.
- 06
Is the outcome fair?
Results meet a stated fairness test on a stated population.
What each result shows — and what it does not.
Every entry links to its source. Where a result comes from research rather than a customer deployment, it says so.
Helix
Serving latency and valid structured output
- What was measured
- HELIX v1.7 compared with stock Ollama using the vLLM project’s GuideLLM harness on the same AMD EPYC 9254 machine, 32K context, at one and four concurrent requests.
- What it does not establish
- One model, one machine and one comparison system. It does not predict results on other hardware, models or serving stacks. The confidence features in v1.8 were added after these runs.
Helix
A 120-billion-parameter model at interactive speed without a GPU
- What was measured
- GPT-OSS 120B throughput on three AMD Zen5 EPYC processors, CPU only, three-run average, run by AMD engineering in September 2026.
- What it does not establish
- Throughput only — it says nothing about answer quality. Results depend heavily on memory configuration, as the published table shows.
Compile-Time Inference
A continually updated decision policy can outperform a frozen one
- What was measured
- A study on a public procurement event log (BPI Challenge 2019), including a control that scrambles labels to test whether the gain is real.
- What it does not establish
- One dataset and one task. It does not show the same gain on your data, and it is a research result rather than a product guarantee.
Compile-Time Inference
When a learned decision policy beats a fixed rule — and when it does not
- What was measured
- A randomised trial and sensitivity analysis identifying the conditions under which a learned policy outperforms a fixed business rule.
- What it does not establish
- The finding includes cases where a fixed rule is the better choice. It is specific to the studied setting.
Compile-Time Inference
The production engine matches its reference implementation
- What was measured
- Zero mismatched selections in 10,000 comparisons between the production engine and the reference implementation.
- What it does not establish
- Shows the two implementations agree with each other. It does not show that either selects the correct decision.
Research paper
Published method: decision policies
- What was measured
- Verificate EAV-DT, arXiv 2606.29280 (June 2026).
- What it does not establish
- A preprint. Read the paper’s own limitations section; results are specific to its workloads.
Research paper
Published method: detecting likely hallucination
- What was measured
- Hallucination prevention (CORTEX), arXiv 2602.17691 (January 2026).
- What it does not establish
- A preprint. Detection rates depend on the model and task, and no detector catches everything.
Gate
Gate caught planted defects that a model’s self-review missed
- What was measured
- Six deliberately planted defects: a frontier model reviewing its own work caught none; Gate caught all six.
- What it does not establish
- A small, adversarial test set built to probe a specific weakness. It is not a general defect-detection rate.
Customer evidence is product-specific. Helix is deployed at UNSW; that does not mean UNSW uses Compile-Time Inference or Gate.
The result that matters is the one on your workload.
A short, well-designed test tells you more than any published figure.
- 01
Write the questions first
Include routine, difficult and unanswerable questions, with reference answers, before you see any system’s output.
- 02
Compare fairly
Test against your current approach configured as well as you can configure it, on the same questions.
- 03
Measure each question separately
Report citation validity, claim support, correctness, repeatability, coverage and refusal on their own lines.
- 04
Count the whole cost
Include preparation, updates, storage, serving and review time — not only the price of a model response.
