Evaluating AI in games needs a trust layer, not another leaderboard

Evaluating AI in games has become one of the harder problems in applied AI, and the industry’s current answer, public leaderboards and automated scores, is failing at exactly the moment the stakes are rising. Game AI is shipping into millions of players’ hands, yet the benchmarks meant to measure it are both gameable and, increasingly, untrusted. The missing piece is not a better score. It is a verifiable trust layer beneath human judgement: proof that a real, unique, qualified person made each evaluation. That is the infrastructure Ontology builds.

Why evaluating AI in games breaks on public scores

The capability is real and advancing. Google DeepMind’s SIMA 2 plays, reasons and self-improves across commercial 3D games (deepmind.google), and NVIDIA ACE companions now ship inside real titles (krafton.com). Measuring whether any of it is good is where evaluation breaks. The most popular human-preference arena has been shown to be gameable, with private mass-testing and large relative gains from fitting the test distribution (arxiv.org). On agentic game benchmarks such as BALROG, current models manage only the easiest environments (arxiv.org). Automated metrics and public leaderboards both overfit. The correction is independent, verified human judgement.

Games are a hard case for a specific reason. The qualities that decide whether game AI is working, whether a companion is enjoyable, whether an NPC holds character across an hour of play, whether an agent pursued the objective or only the metric, are long-horizon and partly subjective. There is no single correct answer to compare against, so the ground truth has to be assembled from human judgement. That makes the reliability of the judges, rather than the sophistication of the scoring, the binding constraint. A benchmark is never better than the people whose opinions it encodes.

The trust problem is now a procurement problem

At the same time, buying human evaluation has become a reputational and legal exposure. Meta took a 49 percent stake in Scale AI in June 2025, and rival labs moved to pull back over conflict of interest (cnbc.com). Scale reportedly left confidential client documents publicly accessible (cvat.ai). Mercor was caught in a supply-chain breach in early 2026 (techcrunch.com). Add the shift to verified expert data as clean web data runs out, with under 5 percent of remaining web content meeting frontier quality standards (signalfire.com), and the market’s message is clear. Labs need evaluation they can trust, from a neutral provider that does not sit inside a competitor and does not hold their raw data.

Evaluating AI in games is therefore a procurement decision as much as a technical one, and that has changed the questions asked in a vendor review. Throughput and price are now the easy part. The questions that decide the contract are closer to these: can you prove your evaluators are distinct people rather than one person with twelve accounts; can you show that a given evaluator has been consistent over months rather than on the day they were onboarded; can you evidence their qualifications without handing us personal data we then have to protect; and does any of our material have to leave our environment for the work to happen. Most vendors answer those with assurances, and assurances are exactly what the last two years have devalued.

What a trust layer has to prove

Evaluating AI in games to a standard a lab can defend comes down to four properties, each of which turns human evaluation from an assurance into a record. Each rests on an open standard rather than a proprietary claim.

Uniqueness. Each evaluator is a distinct human, resistant to duplication and sybil farming. Decentralised identifiers (w3.org) give every evaluator a cryptographically controlled identifier that a buyer can rely on without a central registry to trust.

Longitudinal consistency. A track record, not a snapshot at onboarding. Signed, timestamped verifiable credentials (w3.org) accumulate into a history that shows how an evaluator has actually performed over time.

Qualification without exposure. Selective disclosure lets an evaluator prove the specific claim a buyer needs, domain expertise, hours in a genre, agreement with the pool, without revealing identity or the underlying account.

Revocability and portability. Status has to be checkable and withdrawable. Bitstring Status List (w3.org) lets a verifier check standing without contacting the issuer, and the credential belongs to the evaluator rather than to a vendor database.

None of these are Ontology inventions, and that is the point. A buyer who depends on open standards is not depending on us.

What the trust layer provides

Ontology provides the verifiable identity and reputation infrastructure that turns scattered human judgement into a trusted, auditable signal. Decentralised identity (ONT ID) proves each evaluator is a unique, real human. Verifiable credentials and selective disclosure let evaluators prove their qualifications and track record without exposing who they are. Portable, on-chain reputation gives buyers provenance, proof that a known, consistent human produced each judgement, carried across platforms rather than locked inside one vendor’s database. zkTLS verification confirms authenticity without moving raw data. This is enterprise-grade trust infrastructure for the human side of AI evaluation.

In practice a buyer receives results that carry their own provenance: which verified evaluator produced each judgement, what their record looked like at the time, and how far they agreed with the rest of the pool, all checkable without a single item of personal data changing hands. Neutrality gets an implementation rather than a promise. We do not sit on a competitor’s cap table, and we do not hold the material the judgement was made on.

How this works in practice

The first deployment is Proof of Judgement, an ONTO Wallet campaign recruiting verified gamers as paid gaming AI evaluators. Players connect the platform accounts they already use, which establishes them as genuinely active rather than freshly created. They claim an ONT ID that proves uniqueness. Quality is then measured on the work itself, and the reputation that results is carried in the evaluator’s own wallet. Buyers see a ranked, verified pool. They do not see the people in it.

That order, verify then measure then route, is deliberately demand-led. Evaluation capacity is stood up against real buyer requirements rather than built on spec, which is also how the quality bar stays meaningful once volume arrives. It is what gives evaluating AI in games a defensible ground truth rather than another score.

Gaming first, and what comes after

The first application is gaming AI, because verified players are an ideal, low-risk proving ground and the demand is immediate. The individual side of this, how a player becomes a verified, paid evaluator, is set out in the companion ONTO Wallet piece, Proof of Judgement. The infrastructure underneath is the same trust layer that will serve AI evaluation well beyond games. A user-owned data economy needs a trust layer. Evaluating AI in games is simply where that trust becomes measurable first, and where verified humanity stops being a promise and becomes a record.

Questions AI teams ask

Do we have to move our data? No. Verification confirms authenticity without relocating raw material, and what you receive back is a judgement with provenance attached, not a file of personal data you then have to protect.

How is this different from a vetted expert marketplace? Vetting happens once, at the door, and the result lives in the vendor’s database. Verifiable credentials make the record continuous, independently checkable, and portable, so it survives the vendor relationship rather than depending on it.

What volume can you service? The pool is verified rather than bulk, and it starts focused. For high-stakes evaluation, where a small number of trusted judgements beats a large number of anonymous ones, that is the useful shape. We are not competing on raw volume today and we will not pretend otherwise.

Which standards does this rest on? W3C decentralised identifiers, the verifiable credentials data model, and Bitstring Status List for revocation, with ONT ID and ONTO Wallet as the implementation.

If evaluating AI in games is on your roadmap and provenance has moved onto your vendor checklist, our AI teams page sets out how to work with the verified pool: onto.app/en/for-ai-teams.