# STACK B — LOW-TOKEN ACQUISITION AND RE-ENTRY PROTOCOL

Version: `0.1.1-STACK-B-LOW-TOKEN-PROTOCOL-SEED`  
Status: `DEFERRED EXPERIMENT SEED / CANDIDATE / NON_RUNTIME / NO_AUTHORITY_EFFECT`

## 1. Question

Which compact Stack B surface best preserves binding, recall, error detection, and lawful re-entry under a constrained context budget?

Prompt byte count is recorded but is not the success criterion.

## 2. Compared surfaces

| Surface | Resident material | Intended test |
|---|---|---|
| `TABLE_FULL` | PRE, complete SARTS table, topology and laws | high-information control |
| `PROPOSITIONS_9` | PRE plus the exact nine propositions | semantic compression through predication |
| `COMPACT_SARTS` | nine carrier-tagged `S1|A|R|T|S2` rows plus minimal laws | explicit compact structure |
| `LABELS_ONLY` | `BCDEFGHIK` plus regime labels | lower-bound/known-stack re-entry control |

Every surface must be generated from the frozen fixture or checked against its hash. A trial may not silently paraphrase the nine propositions.

## 3. Isolation

Each run uses a fresh session with no prior Stack B, AEGIS, or CARCER content. Record:

```text
run_id
model_and_version
date
effective_context_or_budget
surface_id
surface_utf8_bytes
surface_token_count_if_available
task_prompt_hash
response_hash
```

Run at least three repetitions per surface and budget condition. Randomized surface order is preferred. Do not treat repetitions inside one continuing conversation as independent.

## 4. Acquisition probes

After exposure, request without supplying the answer:

1. reconstruct all nine carriers and five SARTS roles;
2. explain why S/K is visited twice but represented by one wheel;
3. identify the exact condition for categorical `SER`;
4. distinguish the 9 Stack B diagonals from the 36 Figura Tertia chambers;
5. identify the provenance of the nine propositions;
6. explain why A×T-derived virtues/vices are excluded.

## 5. Re-entry probes

After an unrelated distractor segment within the chosen budget:

1. reconstruct the E record;
2. detect `A_F` inserted into E;
3. reject a raw question in the R slot;
4. preserve nine output positions without invented proposition padding;
5. choose categorical, contracted, suspended, or rejected state from supplied evidence;
6. name the missing ancestry stage in an incomplete RQ record.

## 6. Scoring

| Metric | Range | Meaning |
|---|---:|---|
| carrier/role reconstruction | 0–45 | one point for every correct carrier-bound SARTS value |
| exact proposition recall | 0–9 | exact or predeclared normalized match |
| topology recall | 0–3 | four wheels, S revisit, six incidences |
| drift detection | 0–4 | carrier, RQ, padding, AEGIS violations |
| state decision | 0–4 | categorical, contracted, suspended, rejected |
| provenance/boundary | 0–4 | authorship, AEGIS, virtue/vice, Stack G |
| unsupported invention | negative | subtract one per invented binding or proposition |

Record ambiguous responses separately. Do not force binary success when the answer is underdetermined.

## 7. Success and comparison

A compact surface is behaviorally acceptable only if it preserves the binding and safety scores of `TABLE_FULL` within a predeclared tolerance while reducing resident cost. A lower token count with increased carrier drift, overpredication, or fabricated padding is a failure.

## 8. Scope transfer and execution boundary

Execution ownership has moved to the Backlog request:

```text
M2M-DEPLOY-03
M2M REQUEST — STACK B LOW-TOKEN CO-BOOTSTRAP AND RE-ENTRY BENCHMARK.md
```

That request expands this seed into a 24-run matrix and adds the simultaneous assistant-bootstrap/user-tutorial dependency. The experiment is no longer a completion gate of `M2M-STACK-B-01`.

Results still require genuinely isolated model sessions and recorded model/budget metadata. Static self-inspection in the authoring session is not an empirical result and must not be reported as one. Backlog placement does not activate execution.
