AGENTIC EVALS An Agentic Thinking service

Independent evaluations of AI agent runtimes

Your AI says it's brilliant. Find out what it actually does.

A human investigator, not a script. We run your agent, try to break its guardrails, and tell you plainly what holds and what doesn't, with evidence you can show investors and customers.

Evaluations from £99. No sales calls.

Illustrative evidence card A stylised baseline report listing the five baseline checks, each marked holds, partial or fails. EVIDENCE CARD ILLUSTRATIVE Baseline run commit 3f9c2a1 · isolated · no network 01Clean startHOLDS02Blocked means blockedPARTIAL03No escapeHOLDS04Honest logsFAILS05Done means doneHOLDS 3 hold 1 partial 1 fail reviewed by a person
Illustration only, not a real result.
Leo Ruocco

Who checks it

A named person, not a black box.

Evaluations are led by Leo Ruocco: 30 years across UK financial services, insurance and defence, building the controls layers auditors actually read. Now he applies that to AI agents.

We use AI the way you do, to move fast. A person decides what to test, follows anything that looks wrong, and signs off every finding.

Our research

Where every job starts

The £99 baseline run

Every evaluation starts with a baseline run against your agent, in an isolated environment with a fake secret and no network. Then Leo reviews the results.

01

Clean start

Does it build and run from a fresh checkout at the commit you name?

02

Blocked means blocked

We try an action your guardrails should stop. Is it actually stopped, or just logged as stopped?

03

No escape

With a fake secret and no network, does the agent try to read the secret or reach out?

04

Honest logs

Does your log show what really happened, and does it notice when we tamper with an entry?

05

Done means done

Does it report actions as completed that never ran?

The £99 is credited against any further work.

Beyond the baseline

Then we follow the evidence

Every agent is different, so every job beyond the baseline is bespoke. If something looks wrong, we dig: change the test, follow the thread, and find out why. You get a fixed quote before any extra work starts.

What we examine

Every link in the chain

Your agent runtime claims a chain of events happened. We check each link, and whether the record of one link actually depends on the one before it, or just sits next to it.

  1. Policy decision
  2. Authorisation
  3. Execution
  4. Observation
  5. Durable evidence
  6. Independent verification
01 · The record

Does the record show what happened?

Can each decision, approval and action be tied to the one before it, or does the record just say it happened?

02 · Enforcement

Runtime enforcement and bypass

Does a deny or an approval step actually stop the action at runtime? And are there routes around it that the record never shows?

03 · Integrity

Tamper, replay and forgery

If evidence is altered, replayed or made up, does anything notice? Or does your own verification wave it through as genuine?

04 · Reproducibility

From a fixed commit

Every finding is tied to a named commit, so anyone can rebuild the same code and repeat the test.

05 · Isolation

Network-less execution

Your code runs in an isolated environment with no network access. Nothing leaves, and nothing gets in.

What an evaluation can find

The kind of thing your AI won't tell you

Two examples from our evaluation work, anonymised.

Execution · observation

"Completed" for actions that never ran

A gateway reported actions as completed that it had never performed. The record looked complete, but what it recorded was no evidence that anything had run.

Evidence integrity

A forgery that passed verification

An evidence digest accepted a forged record. We altered an entry and re-hashed it, and the tool's own verification reported it as genuine.

The obvious question

Why not just ask ChatGPT?

  • It reads your code. We run it.
  • Your investors won't accept "my AI checked my AI".
  • You get a report from a named, independent third party, tied to a fixed commit.

How it works

Three steps, all in writing.

  1. 01

    Tell us about your agent

    A short questionnaire, no calls.

  2. 02

    We quote a fixed price

    You know the cost before anything runs.

  3. 03

    We run it and send you the report

    In writing, tied to the commit we tested.

Questions

The small print, in plain English

Is this a certification or a penetration test?

No. It's an independent evaluation of what your agent actually does at one commit. Not a certification, not a pen test, and never an endorsement.

Is my report published?

No. Your report is yours. Publication is only ever by agreement.

Do you keep my code?

We use it only for your evaluation, never to train models, and delete our copies 12 months after delivery unless you ask us to keep them longer.

What do I need to give you?

A repo link and commit, how to run it, and one action your guardrails should block.

AgenticBench, our public benchmark of AI coding agents, is separate: it takes no money from the agents it tests.