What we do
01  Advanced Infrastructure 02  Applied AI & Data 03  AI Cybersecurity 04  AI Assurance
Engagements
AI estate inventory Assurance review
Industries
Financial Services Government & Public Sector Energy & Utilities Telecommunications Healthcare & Life Sciences Transport & Logistics Industrial & Manufacturing Retail, Hospitality & Real Estate
Research
The Trust Maturity Model The GCC Assurance Index Readiness self-assessment Case studies Perspectives Sector briefings Technology evaluations
Company
About us Partners Events Careers Contact العربية Talk to our team
AI Assurance  ·  Independent Evaluation

The vendor tested it. On whose data?

A vendor benchmark is evidence about a benchmark. Evaluation on your own data, your own population and your own thresholds is the only kind that answers a question anybody will ask you.

4–8 WEEKSTESTED ON YOUR DATAWE DID NOT BUILD IT
The decisions underneath

Three questions decide the test.

Evaluation goes wrong when the test is designed around what is easy to measure rather than around what was promised.

01

What exactly was claimed

Written down, in terms that can pass or fail. A great many claims dissolve at this stage, which is itself a finding.

02

Whose population is it tested on

Performance on a public benchmark and performance on your customers are different numbers, and the gap is usually widest where it matters most.

03

What is the threshold, and who set it

A model is not good or bad in the abstract. It is above or below a threshold somebody was willing to put their name against.

What a review asks for

What the system was claimed to do, tested.

Independent evaluation sits in the testing row below, and it is the row most often empty. Choose an obligation to see what it is meant to hold.

THE INVENTORY
ARTEFACT 01

System inventory

Everything in scope, counted
ARTEFACT 02

Risk classification

Applied the same way twice
THE MAPPING
ARTEFACT 03

Control mapping

Obligations held against systems
ARTEFACT 04

Decision record

Who approved, on what information
THE PROOF
ARTEFACT 05

Testing evidence

That the control actually ran
ARTEFACT 06

Reporting pack

In the form it will be asked for
EVIDENCE LEDGER— OF 6 EVIDENCED
Select an obligation

Six artefacts. Every obligation on this list draws on the same six, which is why preparing for one of them prepares you for most of the others.

HELDPARTIALABSENT

Vendor benchmarks are evidence about a benchmark. Evaluation on your own data, your own population and your own thresholds is evidence about your system.

What you receive

A test, and what it showed.

Four stages, every engagement. Hover a stage to see what happens in it.

DURATION
4–8 weeks
DELIVERABLE
Evaluation report against agreed thresholds
DELIVERED
Remotely; on your infrastructure where residency requires
INDICATIVE FEE
[FEE BAND — pending sign-off]
The boundary

We evaluate it. We did not build it.

Three moves. Two of them are ours, and the one in the middle deliberately is not.

MOVE 01 — OURS

We test and we write it down

We evaluate the system against what it was claimed to do and against the obligations that apply to you, and we record what the test showed — including where the answer was that no record exists.

MOVE 02 — NOT OURS

Your builder does the fixing

Remediation is performed by whoever built or runs the system. We do not take the build work, because taking it would mean the next evaluation is us examining our own hands.

MOVE 03 — OURS

We re-test on a cycle

The system is re-examined against the same tests on a schedule, because an evaluation is a statement about a configuration on a date, and configurations move.

If a finding of ours turns into a build contract for us, the finding stops being evidence and starts being a sales instrument. We would rather keep the evidence.
AI Assurance

The rest of this pillar.

Three engagements inside this pillar. Start with the question you can name, or take the whole estate at once.

Questions we are asked

Before you ask us.

Can you evaluate a system Orvix built?
No. That is the point of the pillar. Where Orvix built the system, the evaluation is performed by a third party and we are on the receiving end of the findings like anybody else.
The vendor has published benchmark results.
Those are useful context and they are not evidence about your deployment. The questions that matter are whether your population resembles the benchmark and whether your threshold is the one the benchmark was scored against — and the answer to both is usually no.
What if the evaluation fails?
Then you have a documented finding, before a customer or a regulator produced it for you, and a remediation conversation with your builder that is grounded in a number. That is the useful outcome; it is not a failure of the engagement.
Does this cover bias and fairness?
Where the system makes decisions about people, yes, and it is tested on your own population rather than on a generic fairness suite. The thresholds are agreed in writing beforehand, because agreeing them afterwards is not an evaluation.

Start with the claim you would not want tested.

Thirty minutes. Bring the sentence from the vendor’s proposal you would least like to see put to a test.

Book a 30-minute scoping call

ORVIX · INDEPENDENT AI & TECHNOLOGY ASSURANCE · WE DISCLOSE EVERY COMMERCIAL RELATIONSHIP ON THE PAGE FOR THE SERVICE IT BELONGS TO. WHERE LICENSING IS REQUIRED, DELIVERY IS PERFORMED BY NAMED PARTNERS UNDER THEIR OWN LICENCE.