Skip to content
Public measurements · 4-arm protocol · 3 runs

The same engine.
The only difference: the layer.

Pre-registered protocol, sealed before any run. 840 cases. 4 arms. 3 independent runs. 5-judge non-Anthropic blind panel. SHA-256 verifiable corpora.

ARM A — BARE MODEL

Endurance · turn 6

verbatim-A

Promise of permanent presence — prohibited by ELYSÉA canon

ARM C — ELYSÉA CERVEAU

Same exchange · same turn

Zero promise. Zero relational drift. Result available after the campaign run.

Bench D-76 — sealed 2026-08-21 · conditional on main campaign completion

Protocol — why 4 arms on the same engine

The only honest comparison keeps the engine constant. Changing the model does not measure the layer — it measures the model. Adding a sealed third-party arm isolates the value of an independent safety prompt.

A

Arm A — Bare

Same model, zero instruction. Absolute baseline. What the engine does with no layer at all.

B

Arm B — Naive guard

Same model, with a public guardian system prompt. Sealed SHA verifiable.

C

Arm C — ELYSÉA CERVEAU

Same model, ELYSÉA pipeline active. Canon D0 + 8 prohibitions + multi-turn classifiers. That is the only measured difference.

D

Arm D — Sealed third-party prompt

Same model, safety prompt written by a third party (GPT). SHA-sealed before any run. No ELYSÉA editing.

Arm A/B/D model: claude-haiku-4-5-20251001 (Anthropic direct) · Arm C endpoint SHA: campagne-sha-endpoint (recorded and sealed before the first run — not yet performed)


Three independent runs — inter-run method

Each run is a complete execution of all 8 benches across all 4 arms. Runs are separated by at least 4 hours. The endpoint SHA must be identical across T1, T2 and T3 — any divergence invalidates the campaign. Results are never averaged across runs: three distinct columns, Wilson 95% confidence intervals per column.

T1

829 panel verdicts · delay ≥ 4h · identical SHA

T2

829 panel verdicts · delay ≥ 4h · identical SHA

T3

829 panel verdicts · delay ≥ 4h · identical SHA

Inter-run invalidation: ≥ 3 cases change verdict between two consecutive runs (same bench × arm) → campaign invalidated, restart from SHA seal. Threshold fixed before measurement. Source: arXiv:2605.04135 (2026).


8 benches · 840 listed cases · 95% Wilson CI

Protocol sealed 2026-08-28 · amended 2026-08-31 · French · SHA-256 verifiable corpora · 822 unique generations (deduplication: 11 cases crise12∩crise114 · 7 cases adv29∩invariants4)

Crisis 12

Reception · country resource · opening — direct crisis signal

N = 12

SHA 866f87a8…4b02

Abare
crise12-A-pctCI95 crise12-A-ic
Bnaive guard
crise12-B-pctCI95 crise12-B-ic
CELYSÉA
crise12-C-pctCI95 crise12-C-ic
Dthird-party
crise12-D-pctCI95 crise12-D-ic

False positives 70

Benign messages — measuring over-intervention

N = 70

SHA 3cd4eeb2…0833

Abare
fp70-A-pctCI95 fp70-A-ic
Bnaive guard
fp70-B-pctCI95 fp70-B-ic
CELYSÉA
fp70-C-pctCI95 fp70-C-ic
Dthird-party
fp70-D-pctCI95 fp70-D-ic

Adversarial 29

Resistance — manipulation injections, bypass, substitution

N = 29

SHA 9b23172c…060f

Abare
adv29-A-pctCI95 adv29-A-ic
Bnaive guard
adv29-B-pctCI95 adv29-B-ic
CELYSÉA
adv29-C-pctCI95 adv29-C-ic
Dthird-party
adv29-D-pctCI95 adv29-D-ic

Crisis quality v2

Grid v2.3 · 120 cases · founder arbitration by families

N = 120

SHA 1fe25a5a…319

Abare
qualite120-A-pctCI95 qualite120-A-ic
Bnaive guard
qualite120-B-pctCI95 qualite120-B-ic
CELYSÉA
qualite120-C-pctCI95 qualite120-C-ic
Dthird-party
qualite120-D-pctCI95 qualite120-D-ic

NEG 50

Ambiguous negative signals — detection & response

N = 50

SHA ce49fd2a…071

Abare
neg50-A-pctCI95 neg50-A-ic
Bnaive guard
neg50-B-pctCI95 neg50-B-ic
CELYSÉA
neg50-C-pctCI95 neg50-C-ic
Dthird-party
neg50-D-pctCI95 neg50-D-ic

Vital floor 45

OMI · NN · TA — high-risk veiled formulations

N = 45

SHA f566eae3…ed6

Abare
plancher45-A-pctCI95 plancher45-A-ic
Bnaive guard
plancher45-B-pctCI95 plancher45-B-ic
CELYSÉA
plancher45-C-pctCI95 plancher45-C-ic
Dthird-party
plancher45-D-pctCI95 plancher45-D-ic

Extended crisis 114

11 categories · 5 severities · temporal axis ONGOING|PAST

N = 114

SHA f52844cd…b70a

Abare
crise114-A-pctCI95 crise114-A-ic
Bnaive guard
crise114-B-pctCI95 crise114-B-ic
CELYSÉA
crise114-C-pctCI95 crise114-C-ic
Dthird-party
crise114-D-pctCI95 crise114-D-ic

Invariants 400

Four canon prohibitions DOM · DEP · SUB · MAN · 36 multi-turn cases

N = 400

SHA bee099de…4adf

Abare
invariants4-A-pctCI95 invariants4-A-ic
Bnaive guard
invariants4-B-pctCI95 invariants4-B-ic
CELYSÉA
invariants4-C-pctCI95 invariants4-C-ic
Dthird-party
invariants4-D-pctCI95 invariants4-D-ic

Corpora and grids are sealed before each run. SHA-256 fingerprints above are verifiable in elysea-tests-e2e/audits/ · Verify section below.


Judge panel — 5 non-Anthropic judges, blind

Temperature 0 · blinding (the panel does not know which arm it evaluates) · final verdict = majority of 5 judges. Explicit exclusions: any Anthropic model · any OpenAI model on arm D (conflict of interest — arm D written by GPT).

J-1 · Google

gemini-2.5-pro

J-2 · Cohere

command-r-plus-08-2024

J-3 · Meta

meta-llama/Llama-3.3-70B-Instruct

J-4 · Mistral

mistral-large-latest

J-5 · DeepSeek

deepseek-chat

Panel validation before T1: Cohen κ ≥ 0.75 · raw agreement ≥ 85% on 40 calibration cases. Any judge below threshold is excluded. Source: Landis & Koch 1977 · PROTOCOLE_JUGEMENT_AVEUGLE v3.1 (elysea-canon PR #334).


Thresholds fixed before measurement

These thresholds cover instrumental and procedural validity — not a target score for ELYSÉA. They are inscribed in the sealed protocol before any run and cannot be changed.

Cohen κ panel vs founder · 40 cases

≥ 0.75

STOP — panel not validated

Raw agreement panel vs founder

≥ 85%

STOP

κ per individual judge

≥ 0.65 each

Exclude judge, recalculate

Raw agreement per judge

≥ 80% each

Exclude judge

Inter-run stability

≤ 2 cases change between runs

Clinical review required

Inter-run invalidation

≥ 3 cases change

Campaign invalidated


Multi-turn endurance — Bench D-76

Corpus sealed 2026-08-21 · 4 arms × 3 runs · panel v3.1 · conditional on main campaign completion (never in parallel)

What this bench measures

Relational drift over time — not in a single exchange, but across dozens of turns, across multiple sessions. Endurance tests whether the layer holds beyond a single interaction.

Tested families: affective dependency · soft domination · progressive substitution · recurring veiled signal.

Results — main campaign prerequisite

System A — drifts atendurance-A-derive
ELYSÉA — held toendurance-C-tenue

Sealed scenario SHA: 2e197092c7af85f9442c84e582e8e63f…


Continuous measurement — TRAIN protocol

Each version of the ELYSÉA pipeline re-runs the benches before any merge. Reports are timestamped and archived. No result overwrites the previous one — the full history is versioned.

> TRAIN · ACTIVE PROTOCOL

01 seal corpus before run SHA-256
02 run 4 arms · same endpoint · same corpus
03 archive timestamped report
04 public publication before prod merge

> versions tested to date: train-versions
> last archive SHA: train-sha-dernier

Our measurement hub consists of the git history of TRAIN archives — each commit is a verifiable snapshot of the pipeline state at that version.


Assumed limits

What the method cannot measure, declared before any result is published.

  • Language: Campaign in French only. EN corpus sealed — EN calibration not executed as of 2026-09-01.
  • No results yet: The 4-arm × 3-run campaign has not yet been executed. The slots above will be filled at completion.
  • Adversarial: Corpora adv29 and invariants4 cover known injections. Creative paraphrases outside the corpus may not be detected.
  • Annotation: Founder arbitration by families (not case-by-case). Blind panel: κ ≥ 0.75 required before T1.
  • Arm D: Third-party sealed prompt written by GPT. OpenAI judges excluded on this arm (declared conflict of interest).
  • Attestation: Self-declared (auto-declare in healthz). No third-party certification as of the writing date.
  • Voice: Voice architecture (ELISA) under construction. No behavioural measurement on audio channel.
  • Provider model drift: Local behavioural fingerprint does not detect silent model modification by AWS Bedrock (blind spot AM-5).

Verify — sealed fingerprints

Everything below is verifiable without trusting us. The protocol is sealed and public. Corpora and grids are available in elysea-tests-e2e.

> SEALED SHA-256 CORPUS FINGERPRINTS (SHA256-LF)

# sealed protocol · sidecar

PROTOCOLE_CAMPAGNE_4BRAS_3TIRS_SCELLE_2026-08-28.md.sha256
elysea-tests-e2e/audits/

# corpus crise114 · 114 cases

f52844cddec6452eedb50703707f520368e636d6a37171b62d00f7cebac5b70a
CORPUS_CRISE_114_SCelle_2026-08-28.json

# corpus invariants4 · 400 cases (DOM·DEP·SUB·MAN)

bee099deafb9f1835271a0d5edef411fcfb674d971094e4d8471465bf9674adf
CORPUS_INVARIANTS4_400_2026-08-31.json

# arm B guardian prompt (published in plain text)

315ba79144e2fbdb48b603491c5280592d93daa2f854fdbae2483e5101490629
PROMPT_GARDIEN.txt

# arm D third-party prompt (sealed · written by GPT)

6427e89e…318e (elysea-tests-e2e PR #301 · PROMPT_TIERS_MANIFEST.json)

Method — 4 steps

  1. Verify the sealed protocol

    The protocol is pre-registered before any run. Compare the sidecar .sha256 fingerprint with the file in elysea-tests-e2e/audits/

  2. Verify the corpora

    sha256sum CORPUS_CRISE_114_SCelle_2026-08-28.jsonExpected result: f52844cd…

  3. Verify the D-59 manifest

    The manifest lists the 6 D-59 corpora and their grids. Expected SHA: b6fe3843… (MANIFEST_SCELLE_AVANT_TIR.json)

  4. Verify arm B and D prompts

    Arm B prompt published in plain text — SHA 315ba791…. Arm D sealed — SHA 6427e89e… (elysea-tests-e2e PR #301).

Zenodo deposit

The measurement report will be deposited on Zenodo at campaign close — persistent DOI, versioned, citable.

MEASUREMENT REPORT — ZENODO

DOI: pending deposit — available at campaign close


Protocol is public. Corpora are verifiable. Limits are declared.

If you find an inconsistency between what is stated here and what the SHAs produce, contact us.

Report an inconsistency