# Smallsome MicroRISC: uRISC-T1 Architecture, ISA, Compiler and Formats

As of: 10 October 2026
Theoretical specification: Version 1.0, document revision 4

Branding note: The public brand is **Smallsome MicroRISC**, with the technical
spelling `microrisc`. The normative ISA/model remains `uRISC-T1`, and the
compiler target remains `urisc-t1`. The public entry is
https://microrisc.smallsome.com/, forwarding to the canonical project home
https://smallsome.com/microrisc/. The source language is **Smallsome 3.0**, based
on Cymple 2.0. Its source modes are **Smallsome Code** (ASCII syntax) and
**Smallsome Symbols** (Unicode-symbol syntax). The format family
is **Smallsome Formats**; individual format names and source filenames remain
unchanged.

Document revision 2 applied English text and confirmed Smallsome branding.
Revision 3 aligned the language references with the shared 3.0 specification.
Revision 4 corrects the num/word boundary, ownership, cancellation and language
manifest contracts. ISA version, encodings and hardware arithmetic are unchanged.
All compiler, runtime and lab contracts remain theoretical; no implementation
or physical validation is commissioned by this document.

## 1. Purpose and Status

This document consolidates the current design of the uRISC CPU family.
It distinguishes binding target decisions, justified target proposals, and the
historical state of the existing uRISC-Lab-v4.1 emulator.

The binding theory now also covers all opcodes, their encoding,
compiler/ABI rules, integration of all eight in-house format families, and
the contract for a future 8x8 lab app with a Scrapbook.
No app, RTL, or hardware prototype is being built at this stage.
Fully specified does not mean implemented or physically proven.

The central idea is not a particularly large single processor. Many small,
deterministic cores are scheduled into a static task and data flow by the
compiler itself. Specialized hardware is limited to time-critical
transport, particularly display scanout, audio modulation, DMA,
RAM access, and USB at the bit level. Audio, video, graphics, and
application algorithms remain programmable.

Current target scope (decided): exactly one CPUlet with 64 cores and
256 MiB of shared RAM. All decisions in this document refer to
this target system. Systems with multiple CPUlets remain open and are collected
as future work in section 23.

Terms used in this document:

- **decided**: target decision for the hardware architecture.
- **normative**: binding behavior of the theoretical T1 model.
- **target proposal**: architecturally preferred, but not yet encoded as RTL.
- **emulator state**: behavior of the existing v4.1 model; not automatically
  a decision for future hardware.

Reading order:

- Sections 2..8: computer structure and cores.
- Sections 9..11: complete binary ISA.
- Sections 12..20 and 27: memory, contact, events, and devices.
- Section 28: compiler, ABI, and runtime.
- Section 29: all in-house formats with their actual dependencies.
- Sections 30..31: lab/Scrapbook, evidence, and source baseline.
- Sections 21..25: historical hardware guidance and future work, not
  prerequisites for the theoretical lab.

The historical reference is located at:

`Tools/Diverse Tools/murisc/murisc_v41_mod_webm_opcode_audit.zip`

## 2. Binding Guidelines

1. A core is small, deterministic, and fully in-order.
2. Normal instructions must not cause wait cycles. Fixed, documented
   execution times such as the branch penalty are not wait cycles.
3. The result of an instruction that writes a register must be usable by the
   immediately following instruction.
4. External latencies are not pulled into the instruction pipeline.
5. Only the explicit WAIT family may put a running task to sleep.
   DONE ends it; FAULT and dispatcher halt are separate control states.
6. Local memory is an explicit scratchpad, not a cache.
7. Shared RAM, display, USB, disk I/O, and other peripherals use
   the same contact contract.
8. A uniform contact means a uniform protocol, not a
   single electrically shared bus.
9. There is no cache coherence. Data ownership and synchronization are managed by
   the compiler, dispatcher, messages, and explicit transfers.
10. The ISA remains independent of the number of available cores.
11. No codec-, audio-, or graphics-specific opcodes are introduced.
12. Future larger systems are built from identical 64-core units
    (future work, section 23). Current decisions must not
    prevent this scaling.
13. Each local 1-KiB slot has exactly one owner per cycle.
14. Response data for a transaction initiated by a core lands
    exclusively in the local SRAM reserved for it. Autonomous bus masters such as
    USB, display, and audio DMA may use shared RAM directly.
    Every completion, event, and error reaches the responsible
    core as a message. There is no second notification path.
15. No additional protocol or graphics bridges are intended.
    DRAM, power supply, clock source, connectors, and the
    protection, level-shifting, and load components required for each
    connection are unavoidable.

## 3. Structure

### 3.1 Core

The core is the smallest programmable computing unit.

### 3.2 CPUlet

A CPUlet consists of exactly 64 cores and a local contact router.
The target system consists of exactly one CPUlet.

```text
1 CPUlet = 8 groups x 8 cores = 64 cores
```

The term CPUlet initially denotes a logical and physical
functional unit on a chip. It does not imply a future division into
separately manufactured chiplets.

Internal structure (target proposal):

```text
8 cores --+
   ...    +-- Group arbiter (8:1 / 1:8, round robin within the class) --+
8 cores --+                                                           |
                                                                       |
     8 group ports + RAM/peripherals + CPUlet-SRAM ---------------------+
     = 10x10 crossbar, 64 bits
```

- A full 64-port crossbar is ruled out because its wiring grows quadratically.
- A group port provides 4 GB/s of raw bandwidth and at most 3.2 GB/s of
  full-data payload bandwidth for eight cores: when shared fairly, at most
  400 MB/s or 1.6 bytes per RUN cycle and core.
  The shared RAM port and group/router contention jointly determine the
  bottleneck; the topology alone does not prove freedom from bottlenecks.
- Messages between cores in the same group do not use the router.
- The group is also the unit for clock gating and shutdown.
- Four mesh ports for multi-CPUlet systems belong to future work
  (section 23) and are not present in the target system.

## 4. Target System

The decided configuration is:

```text
Cores:                64 (8 groups x 8)
CPUlets:               1
System cores:          4 permanently reserved
Application cores:    60
Clock levels:          0 / 50 / 250 MHz
Local memory:         12 KiB per core
CPUlet-SRAM:          0 bytes in the T1 standard profile, address window reserved
Shared RAM:           256 MiB
Graphics output:      up to 1280 x 720 at 60 Hz via DVI-D
Audio output:         up to 7.1 via PDM, alternatively I2S/TDM
Input devices:        3 x USB 1.1 Full-Speed with integrated LS/FS PHYs
Mass storage:         1 x USB 2.0 High-Speed
```

The four system cores handle input, output, USB, the dispatcher, and other
system tasks. The remaining 60 cores are available to applications and media.
The exact internal distribution of the four system tasks remains a matter
for the compiler and operating system.

Minimum board population: uRISC chip, one LPDDR4 component
(section 15.2), power supply, clock source, connectors, and the
required passive protection and termination components. USB port power and
an analog audio output capable of driving its load may require additional
load switches, buffers, or amplifiers. DRAM in the same package (SiP) is a
possible future variant.

## 5. Clock and Energy Model

The previously discussed 750-MHz level is removed. The target hardware has three
states:

| State | Clock | Meaning |
|---|---:|---|
| `OFF` | 0 MHz | shut down or only necessary retention |
| `IDLE` | 50 MHz | light real-time and standby work |
| `RUN` | 250 MHz | full computing performance |

Omitting 750 MHz is a deliberate architectural decision:

- A cycle at 250 MHz lasts 4 ns rather than 1.33 ns.
- Single-cycle multiplication, local memory, and complete forwarding are
  substantially more realistic.
- Supply voltage, clock tree, area, and power dissipation decrease.
- The compiler distributes bulk data work across multiple cores instead of
  forcing peak serial performance.

The contact fabric, RAM controllers, and display unit may independently
operate at 500 MHz. Frequency changes and shutdown should be controllable at
least at CPUlet or group level (8 cores); an individual
core must be capable of separate clock gating and being set to `OFF`.

## 6. Computing Requirements and Evidence Status

The previously discussed core counts are planning assumptions. None has
yet been confirmed by a cycle-accurate T1 guest program.
Native Mac measurements in the format matrix demonstrate the listed
decoder runs, not the instruction costs of our ISA.

| Task | Hypothesis for T1 validation |
|---|---|
| RAU | one IDLE core for a specified mono/stereo stream |
| RFXL | one IDLE core for a specified image size and loading deadline |
| PMF0 | four RUN cores for a specified video/bitrate profile |
| 3D in the class of Tomb Raider 1 | about six to eight RUN cores for a bounded scene case |
| System and I/O | four reserved cores; specific system load must be budgeted |
| Office/media load | application-specific; 16/64 cores indicate scale |

RAU requires more than local SRAM for its unchanged native main stereo buffers.
RFXL streams may have serial dependencies.
PMF0 requires reference RAM. Section 29 describes these limits
and the specific form of validation for each format.

64 cores at 250 MHz yield a theoretical 16 GInstr./s at full local throughput;
after reserving the four system cores,
15 GInstr./s remain for applications.
Branches, serial portions, slot planning, contact, and RAM limit
actual performance. Estimating an instruction count from an ARM runtime
without a trace is not proof of T1 performance.

## 7. Compiler and Execution Model

The compiler is a central part of the architecture. During compilation it should
already:

- split loops and independent data regions,
- distribute tasks across cores or core groups,
- schedule local 1-KiB slots and define their ownership for each phase,
- allocate RAM regions and resources,
- schedule communication and completion events,
- check real-time budgets in cycles, including branch penalties,
- define data ownership so that no cache coherence is required.

The hardware dispatcher starts the prepared tasks. It should not be replaced by
complex dynamic out-of-order or migration logic.

Serial portions remain possible. USB state machines, filesystem logic,
linked structures, and heavily branching control do not have to be artificially
parallelized. 250 MHz is considered sufficient for these control tasks;
DMA and specialized transport handle bulk data.
Light system tasks such as mouse and keyboard run on the system cores and
start suitable routines on other cores as needed.

## 8. Target Core Structure

### 8.1 Registers and State

The v4.1 emulator provides the current basis:

- 12 general-purpose 32-bit registers `R0` through `R11`,
- 4 resource registers `S0` through `S3`,
- 32-bit program counter,
- 64-bit accumulator for `MAC`,
- persistence and active masks for local memory banks,
- timestamps for `WAIT`,
- instruction counter,
- message FIFO with 16 entries.

In the emulator, a message contains a 16-bit sender identifier and a
32-bit value. 16 bits are also sufficient for future multi-CPUlet systems with
up to 4096 cores; the final subdivision of the identifier by group and core
is not separately encoded by CPUlet/group in the historical emulator.
T1 uses coreID=8*y+x; the future extension is not part of T1.

T1 adds kind8/status8; the message is exactly 64 bits.
Core and resource identifiers and buffers are defined in section 27.

The register file contains twelve 32-bit values, three read ports for SEL,
and up to three write capabilities for RECV. Physical banking,
bypass, and port area remain to be determined; they are not a measured
negligible quantity.

### 8.2 Local Memory

Each core has a fixed:

```text
12 slots x 1 KiB = 12 KiB local SRAM
```

Properties:

- guaranteed local and deterministic,
- target: one-cycle access time,
- explicitly managed by program and compiler,
- individually activatable banks,
- persistence mask for `IDLE` and `WAIT`,
- no transparent cache,
- no automatic cache fill or write-back.

Slot ownership (decided):

- Each 1-KiB slot is a separate SRAM bank.
- Per cycle, a slot has exactly one owner: instruction fetch, load/store unit,
  or contact/DMA.
- Code starts at slot 0 by default but may reside in any slot. The
  compiler determines placement during optimization.
- The compiler guarantees that a slot in a given phase is either code, data,
  or a DMA destination.
- A violation causes a `FAULT` rather than a wait cycle
  (section 19).

Four MiB of shared RAM per core is only a capacity rule for external RAM.
It is not installed as physically local SRAM per core and
is not permanently assigned to a core.

### 8.3 Pipeline

A four-stage in-order pipeline is decided:

| Stage | Task |
|---|---|
| IF | instruction fetch from the local code slot |
| ID | decoding, register reads, and forwarding preparation |
| EX | ALU, multiplication, MAC, address calculation, local SRAM access, and branch decision |
| WB | write-back |

Observable contract:

- at most, and normally, one instruction per active cycle,
- no out-of-order execution,
- no speculative program execution that changes state,
- no normal instruction causes a wait cycle,
- the result is available to the immediately following instruction, including after `LD`,
- complete forwarding for register results,
- fixed and documented execution time for all normal instructions.

Branch penalties (decided):

| Instruction | Additional lost issue cycles |
|---|---:|
| `BEQ`/`BNE` not taken | 0 |
| `BEQ`/`BNE` taken | 2 |
| `JMP` | 2 |

The branch decision occurs in EX. Instructions fetched up to that point may
be decoded, but must not execute when the branch is taken; they are
discarded before any state change. An `LD` immediately before a branch
supplies its result through WB-to-EX forwarding. This causes neither a
wait cycle nor a combined SRAM-compare-PC path within the same
cycle.

Fetch/decode errors of a younger instruction are only recorded as pending until
its valid EX acceptance. A discarded wrong path causes no FAULT.
Branches and faults check their own state before any side effect.

Branches with magnitude comparisons are not introduced; `CMP`
and `BNE` serve that purpose. The compiler knows the fixed penalty and can
express short paths without branches using `SEL`.

More area for parallel computing paths, fast multiplication, and
forwarding is accepted. Variable or iterative instruction execution times are
not accepted.

## 9. Instruction Format and ISA Version (Normative)

The theoretical ISA is named `uRISC-T1`, version 1.0. It has exactly
31 primary opcodes. T1 denotes the fully specified theoretical
contract, not an implemented chip or one validated by RTL.
Old v4.1 machine code is not binary compatible. Source programs are
reassembled; there is no automatic legacy execution mode.

An instruction word is 32 bits and is stored little-endian in local SRAM:

```text
31          26 25      22 21      18 17      14 13             0
+-------------+----------+----------+----------+----------------+
| Opcode 6 Bit| A 4 Bit  | B 4 Bit  | C 4 Bit  | Imm 14 Bit    |
+-------------+----------+----------+----------+----------------+
```

`I14` is the signed immediate in bits 13:0; `U14` is the same
bit value without a sign. `I22/U22` use bits 21:0, `U18` bits 17:0.
Register identifiers 0 through 11 denote R0 through R11. R0 is a normal,
writable register. 12 through 15 are invalid for GPR operands.
Resource fields separately address S0 through S3; they are not GPR aliases.

Every undocumented mode, reserved opcode, and set
reserved bit causes `FAULT_ILLEGAL`. Unused fields must be zero.
Invalid register fields are not silently masked.

Addressing is byte-based. The PC is the address of the current instruction;
`nextPC = PC + 4`. A relative branch calculates
`target = nextPC + 4 * signExtend(displacement)` without 32-bit overflow.
PC and branch target must be divisible by four and reside in an enabled,
active code slot. An invalid target causes `FAULT_PC` before
the link register or PC is changed.

## 10. Historical v4.1 Opcode Set (Informative)

The historical `isa.pbi` numbers opcodes starting at 1:

```text
NOP MOVI MOV ADD ADDI SUB AND OR XOR SHL SHR MUL MAC
LD32 ST32 LD16 ST16 BEQ BNE JMP WAIT WAITUNTIL
SEND RECV SIGNAL RWRITE RREAD DONE
```

These are 28 primary opcodes. v4.1 is a functional starting point:
14-bit MOVI, separate load/store opcodes, partially blocking
resource instructions, and a simulation quantum instead of a pipeline.
This meaning applies exclusively to the archived version.
Section 11 is the sole normative T1 ISA.

## 11. Complete Opcode Set (Normative)

### 11.1 General Arithmetic Rules

`u32(x)` denotes the lower 32 bits; `s32(x)` is the same bit value in
two's complement. Normal addition/subtraction and immediate addition
operate modulo 2^32. They produce neither flags nor exceptions on overflow.
There is no flags register. A comparison explicitly produces a mask.

Lane 0 occupies the least significant bits. Width `w` and lane count `n`
satisfy `w*n=32`. Carries, saturation, and shifts operate per lane;
they never extend into an adjacent lane.

The 64-bit accumulator `ACC` is a two's-complement bit vector.
Accumulation is modulo 2^64, without implicit saturation.
All operands are read before the result is written. This also applies
when destination and source are the same register.

### 11.2 Opcode Table

Numbers 16 and 17 remain reserved so that historical LD16/ST16 words
do not accidentally denote a new valid instruction. 0 and 34 through 63
are also reserved: 31 assigned, 33 reserved.

| Decimal / Hex | Instruction | Canonical operands | Effect |
|---|---|---|---|
| 1 / 01 | NOP | none | nextPC only |
| 2 / 02 | MOVI | Ra, I22 | Ra = signExtend(I22) |
| 3 / 03 | MOV | Ra, Rb | Ra = Rb |
| 4 / 04 | ADD | Ra, Rb, Rc, mode | per-lane addition |
| 5 / 05 | ADDI | Ra, Rb, I14 | scalar modulo addition |
| 6 / 06 | SUB | Ra, Rb, Rc, mode | per-lane subtraction |
| 7 / 07 | AND | Ra, Rb, Rc | bitwise AND |
| 8 / 08 | OR | Ra, Rb, Rc | bitwise OR |
| 9 / 09 | XOR | Ra, Rb, Rc | bitwise exclusive OR |
| 10 / 0A | SHL | Ra, Rb, count, mode | per-lane left shift |
| 11 / 0B | SHR | Ra, Rb, count, mode | per-lane right shift |
| 12 / 0C | MUL | Ra, Rb, Rc, mode | half of a scalar 64-bit product |
| 13 / 0D | MAC | Ra, Rb, Rc, mode | write/accumulate/read ACC |
| 14 / 0E | LD | Ra, [Rb + I14], mode | local 8-/16-/32-bit load |
| 15 / 0F | ST | Ra, [Rb + I14], mode | local 8-/16-/32-bit store |
| 18 / 12 | BEQ | Ra, Rb, I14 | branch on bit equality |
| 19 / 13 | BNE | Ra, Rb, I14 | branch on bit inequality |
| 20 / 14 | JMP | form, operand | relative/indirect jump, optional link |
| 21 / 15 | WAIT | form, operand | explicit wait for message/time |
| 22 / 16 | WAITUNTIL | Ra, Rb | explicit wait until absolute time |
| 23 / 17 | SEND | Ra, Rb, Rc | status to Ra, destination from Rb, value from Rc |
| 24 / 18 | RECV | Ra, Rb, Rc | status, value, and metadata from Rx FIFO |
| 25 / 19 | SIGNAL | Ra, Rb | SEND with fixed value 1 |
| 26 / 1A | RWRITE | Ra, Sb, Rc | status, resource, descriptor address |
| 27 / 1B | RREAD | Ra, Sb, Rc | status, resource, descriptor address |
| 28 / 1C | DONE | none | end task and stop core |
| 29 / 1D | CMP | Ra, Rb, Rc, mode | per-lane zero/full mask |
| 30 / 1E | SEL | Ra, Rb, Rc | bitwise selection using old Ra as mask |
| 31 / 1F | PERM | Ra, Rb, Rc, pattern | select four bytes from eight source bytes |
| 32 / 20 | MOVHI | Ra, U18 | replace upper 18 bits |
| 33 / 21 | CLZ | Ra, Rb | count leading zero bits |

`Ra/Rb/Rc` are fields A/B/C, except where the following tables define
a different field assignment. A field not required for the particular
instruction is zero. Modes are bit values, not additional instruction words.

### 11.3 Constants, Moves, and Boolean Operations

- NOP: A=B=C=U14=0.
- MOVI: A is the destination, bits 21:0 are I22; range -2097152 to 2097151.
- MOV: A destination, B source; C=U14=0.
- MOVHI: A destination, B=0, bits 17:0 are U18.
  `Ra = (U18 << 14) | (oldRa & 0x3FFF)`.
- AND/OR/XOR: A destination, B/C sources; U14=0.
- ADDI: A destination, B source, C=0; I14 ranges from -8192 to 8191.

An arbitrary 32-bit constant K is canonically constructed using:

```text
MOVI  Ra, K & 0x3FFF
MOVHI Ra, K >> 14
```

For K in the I22 range, MOVI with the signed value is sufficient.
The assembler may optimize the two-instruction sequence but must not
generate additional undocumented opcodes.

### 11.4 ADD, SUB, and CMP

ADD/SUB: bits U14[1:0] encode lane width:
0=32, 1=16, 2=8, 3=invalid. Bits [3:2] encode
0=modulo, 1=unsigned saturating, 2=signed saturating, 3=invalid.
Bits [13:4] are zero.

Unsigned saturation clamps to [0, 2^w-1], signed saturation to
[-2^(w-1), 2^(w-1)-1]. SUB with unsigned saturation returns zero if
the difference would be negative.

CMP: bits [1:0] are the same width identifier. Bits [4:2] select:

| Value | Comparison per lane |
|---:|---|
| 0 | equal |
| 1 | unequal |
| 2 | signed less than |
| 3 | signed less than/equal |
| 4 | unsigned less than |
| 5 | unsigned less than/equal |

6/7 are invalid; bits [13:5] are zero. A true condition returns
all w bits set, a false condition all zero. Greater than and greater than/equal
are expressed by swapping sources.

SEL has U14=0 and computes
`Ra = (oldRa & Rb) | (~oldRa & Rc)`.
A scalar CMP mask selects a whole word; lane masks select lanes,
arbitrary masks select individual bits. SEL is not a truth-value test.

### 11.5 Shifts and PERM

SHL/SHR use the following bits:

| Bits in U14 | Meaning |
|---|---|
| [1:0] | lane width as in ADD |
| [2] | SHR: 0 logical, 1 arithmetic; SHL: must be 0 |
| [3] | 0 immediate count, 1 register count |
| [8:4] | immediate count, only when [3]=0 |
| [13:9] | zero |

For an immediate, C=0. For a register, C is the count source and [8:4]=0.
All lanes use the same count; its lower log2(w) bits apply.
Immediate values >=w are invalid. Register counts are used modulo w.
Vacated positions are filled with zero for logical shifts and with the old
lane sign for arithmetic SHR. `SAR` is an assembler representation
for SHR with bit [2]=1, not a separate opcode.

PERM: bits [2:0], [5:3], [8:6], [11:9] select result bytes
0 through 3. Index 0..3 selects Rb byte 0..3, 4..7 Rc byte 0..3.
Bit [12]=1 reads Rc as four zero bytes; C must then be 0 and Rc is
not read. Bit [13]=0. Example byte swap of Rb: indices 3,2,1,0,
pattern 0x0053 with the zero source optionally disabled.

### 11.6 MUL and MAC

MUL has exclusively scalar 32-bit sources. Lane MUL is not part of T1.
U14[0]=0 unsigned, 1 signed; [1]=0 lower, 1 upper product half.
[13:2]=0. Signed MUL multiplies two s32 values, unsigned MUL two u32 values.
The lower half is identical for both; the upper half depends on signedness.
There is no automatic rounding or saturation.

MAC: [0] selects unsigned/signed product as in MUL, [2:1] selects:

| Value | Action |
|---:|---|
| 0 | ACC = u64(ACC + product); Ra = low32(ACC_new) |
| 1 | ACC = product; Ra = low32(ACC_new) |
| 2 | Ra = low32(ACC); ACC is retained |
| 3 | Ra = high32(ACC); ACC is retained |

[13:3]=0. For action 2/3, B=C=[0]=0. Action 1 with a
zero-valued register as source sets ACC to zero. This is the canonical
initialization; no hidden reset opcode is introduced.
Immediately successive MAC instructions see the new ACC through forwarding.

CLZ: C=U14=0, Ra is the number of leading zero bits in Rb; CLZ(0)=32.
DIV, MOD, CTZ, LOOP, floating point, and specialized cryptographic opcodes are
not part of T1. Software libraries implement them where permitted by the
target profile. A LOOP proposal is therefore rejected for T1.

### 11.7 Local Memory

For LD/ST, C is a mode field, not a register:

| C | LD | ST |
|---:|---|---|
| 0 | u8 to u32 | lower 8 bits |
| 1 | u16 to u32 | lower 16 bits |
| 2 | u32 | 32 bits |
| 3 | s8 to s32 | invalid |
| 4 | s16 to s32 | invalid |

The effective address is formed as the mathematical sum u32(Rb)+I14.
Negative values, overflow, an end beyond 12288, and missing permissions fault.
16-/32-bit accesses are aligned to 2/4 bytes. An access must not cross a
slot boundary. Byte order is little-endian.
An ST writes only after all checks; on FAULT, SRAM remains unchanged.

The IF code slot and EX data slot in a cycle must be different.
A DMA-owned slot must not be read or written by IF or EX.
Violation: FAULT_SLOT. Code may execute only after transfer completion and
explicit release by the dispatcher (section 27).

### 11.8 Branches and Software Functions

BEQ/BNE: A/B sources, C=0, I14 relative to nextPC.
No flags are read. CMP plus comparison with a known
zero-valued register implements magnitude comparisons.

JMP uses A as the form:

| A | B/C/Imm | Effect |
|---:|---|---|
| 0 | bits 21:0 = I22 | relative jump |
| 1 | B=target register, C=Imm=0 | absolute indirect jump |
| 2 | bits 21:0 = I22 | relative; R11 = nextPC |
| 3 | B=target register, C=Imm=0 | indirect; R11 = nextPC |

Other A values are illegal. For form 3, an old R11 used as the jump source
is read before writing the link. All taken forms cost exactly two
additional issue cycles. `call label` and `ret` are unambiguous
assembler pseudoinstructions for form 2 and form 1 with B=11.
There is neither a hardware stack nor separate CALL/RET opcodes.

### 11.9 Waiting and Task Completion

WAIT and WAITUNTIL form the only explicit wait instruction family.
They complete once; nextPC is then the resume point.
An already present message prevents entry into WAITING.
An event arriving simultaneously must not be lost.

WAIT uses A as the form:

| A | Operand | Effect |
|---:|---|---|
| 0 | B=C=Imm=0 | wait until Rx is not empty |
| 1 | bits 21:0 = U22 | wait until now+U22 system ticks or a message |
| 2 | B/C GPR, Imm=0 | duration = (u32(Rc)<<32) OR u32(Rb); time or message |

WAITUNTIL: A is low32, B high32 of an absolute system-tick deadline,
C=Imm=0. Here too, a message may wake the core early.
Comparisons are modulo 2^64; a time interval must be smaller than 2^63.
Zero duration or an already reached deadline does not wait.
In T1, a system tick is 2 ns (500 MHz), independent of the core clock.
An absolute time is read through the system resource;
no invented host time influences the guest.

A timer wake is latched as a wake reason, creates no artificial
message, and consumes no Rx entry. WAIT observes state;
this is not a second external data or notification channel.
Software that wants to wait exclusively until the deadline checks the time
again after a message wake.

DONE: A=B=C=Imm=0. All older instructions are completed;
younger ones are discarded. With open transfers, occupied Tx, or
reserved completions, DONE causes FAULT_INFLIGHT. Otherwise the
core enters STOPPED. The dispatcher takes over resources and state.
STOPPED is not sleeping until the next message.
Completion generates TASK_DONE through the same contact. A reserved
completion latch retains it until ACK; a new task image may start only
afterwards. This latch is not a normal task Tx entry.

### 11.10 Message Instructions

SEND: A status destination, B destination core, C message value, Imm=0.
Resource/system messages cannot be forged through SEND.
SIGNAL: A status destination, B destination core, C=Imm=0; value is 1.
Success means insertion into the local Tx buffer, not yet
acceptance by the destination.

RECV: A status destination, B value destination, C metadata destination, Imm=0.
All three destination registers must be different. On success, exactly
one FIFO entry is removed:
Rb = value32; Rc = source16 | (kind8<<16) | (eventStatus8<<24).
Ra = OK. If the FIFO is empty, Ra=EMPTY; Rb/Rc remain unchanged.
SEND/SIGNAL may overlap destination and source registers because sources are
read first. No message instruction blocks.

Message formats, statuses, and reserved completion slots are defined once
in section 27. There is no implicit IRQ routine.

### 11.11 Resource Instructions

RREAD/RWRITE: A status destination, B resource identifier 0..3, C GPR containing
a local descriptor address, Imm=0. The descriptor is 16 bytes, divisible by 4,
entirely in a DATA slot, and readable by EX. It is read synchronously;
only then may it be changed.

| Offset | Type | Field |
|---:|---|---|
| 0 | u32 | offset within the resource |
| 4 | u16 | local SRAM address |
| 6 | u16 | length, 1..1024 bytes |
| 8 | u32 | cookie for completion message |
| 12 | u32 | reserved, must be 0 |

The data resides entirely within a single enabled DATA slot.
Descriptor and data must occupy different slots. An RREAD destination
is exclusively owned by the contact until completion, as is an RWRITE source.
IF/EX may continue using other slots.
Hardware atomically validates resource bounds, rights, destination, length,
a free tag, and a free completion slot before acceptance. On rejection,
the data slot remains free and no completion message is generated.

An accepted request returns Ra=OK and exactly one later completion.
Local range/rights errors return a status; invalid
instruction fields or a descriptor slot already owned by DMA
cause FAULT. RREAD never writes to a register later.

### 11.12 Timing and Assembler Contract

All normal instructions have an EX issue window of one
core cycle. The four-stage pipeline has four cycles of latency from IF to WB,
but a throughput of at most one instruction per cycle.
A taken branch discards the two younger windows.
WAIT and DONE drain the pipeline in order. FAULT is precise.

A descriptor access internally requires a 128-bit read operation from its
bank; ordinary LD/ST provide at most 32 bits. This is an explicit
area/bank requirement of the theoretical model. A physical design
requiring multiple cycles here does not satisfy T1 unchanged.

The assembler supports labels, decimal/hex literals, registers,
symbolic mode names, `.code`, `.data`, `.align`, `.byte`, `.word`,
and defined pseudoinstructions:
load32=constant sequence, call/ret=JMP forms, SAR=SHR mode.
It checks all values before encoding and reports address, source line, and
rejection reason. It does not silently expand oversized relative branches;
the compiler explicitly generates load32 plus indirect JMP for them.

Encoding examples for registers R0/R1/R2:

| Assembler | Hex word | Bytes in SRAM |
|---|---|---|
| NOP | 04000000 | 00 00 00 04 |
| MOVI R0, 1 | 08000001 | 01 00 00 08 |
| ADD R0, R1, R2, scalar.wrap | 10048000 | 00 80 04 10 |
| JMP.rel -1 | 50000000 OR 003FFFFF = 503FFFFF | FF FF 3F 50 |

These examples are theoretical encoding requirements, not recorded
assembler test runs.

## 12. Wait-Free Instructions and External Events

Normal arithmetic instructions, local memory instructions, and message operations
do not block the pipeline. External resources physically cannot
guarantee an immediate response. Therefore:

- `RREAD` and `RWRITE` specify a local SRAM address and a length.
- An `RREAD` is sent only if a free transaction tag exists.
  The response slot is therefore already reserved when the request is made.
- Response data lands directly in the specified local slot, never asynchronously
  in a register. Late responses cause no register hazards.
- Completion generates a message in the Rx FIFO.
- `SEND`, `RECV`, `RREAD`, and `RWRITE` immediately return a status in a
  register. A full or empty FIFO and a missing tag do not block the
  instruction.
- Software may then deliberately wait, process another task, or
  retry later.
- Only the explicit WAIT family puts a task to sleep; messages
  and the documented timer conditions may wake it.

This is a change from the v4.1 emulator. There, `RECV`,
`RREAD`, and `RWRITE` reset the program counter when data is missing and
put the core to sleep.

## 13. Messages and Synchronization

### 13.1 Buffers per Core (Decided)

| Structure | Size | Meaning |
|---|---:|---|
| Rx message FIFO | 16 | shared acceptance order |
| normal Rx quota | 8 | USER/system input |
| reserved completion quota | 8 | open transfers and unread completions |
| Tx buffer | 8 | outgoing messages until ACK |
| transaction tags | 8 | RREAD/RWRITE including completion consumption |

The exact quota, reservation, and status semantics are defined in
sections 27.1 and 27.5. Full structures produce a local status;
they do not stop any normal instruction.

### 13.2 Delivery

- `SEND` reports success when there is space in the Tx buffer.
- Each occupied Tx entry carries a unique message identifier until
  completion. ACK and NACK contain this identifier.
- The entry remains occupied until the destination acknowledges with ACK. A retry
  uses the same identifier; the destination must not insert an already accepted
  message into its Rx FIFO a second time.
- If the destination FIFO is full, the destination responds with NACK. Hardware
  automatically retries delivery with fair arbitration.
- Messages from the same sender to the same destination are
  accepted in sending order. There is no global order
  between different senders.
- Messages are rejected at the destination rather than queued in the network.
  They never permanently block links. This rules out protocol deadlock due to
  a full destination FIFO. Logical software deadlock, where tasks wait for
  each other's events, remains possible and must be prevented by the compiler
  or runtime system.

### 13.3 Use

Messages are used for:

- task completion,
- data ownership changes,
- events,
- timers,
- DMA and transaction completion,
- resource readiness,
- error reports (`FAULT`),
- waking a sleeping core.

RAM regions writable by multiple participants should not be synchronized through
implicit cache or lock logic. The preferred sequence is:

1. The compiler or dispatcher assigns a region to a writing task.
2. The task writes its data.
3. A completion event transfers ownership.
4. Subsequent tasks read or take over the region.

Atomic RAM operations are not yet decided. They should be introduced only
if messages and static data ownership are demonstrably insufficient.

## 14. Uniform Contact

### 14.1 Basic Parameters (Decided)

```text
Data width:                  64 bits
Contact clock:               500 MHz
Raw bandwidth:                4 GB/s per link and direction
Maximum transaction:         1024 bytes
Header per packet:           16 bytes
Payload per packet:          at most 64 bytes
Full transfer data direction: 16 x (2+8) = 160 beats = 320 ns
Full-data payload bandwidth: at most 3.2 GB/s before arbitration
Address width:               40 bits
```

Smaller transfers with byte precision are allowed. Final beats
do not become visible as additional payload bytes.
A full 1-KiB transfer corresponds to one slot, but shorter
transfers also reserve the entire affected bank.

### 14.2 Channels

Request and response are separated logically and in terms of credits.
Source, destination, address, tag, length, class, and status reside in the
uniform header. Its complete bit allocation is defined once
in section 27.7. Messages use the request channel, ACK/NACK
the response channel. Pure control responses have a reserved
opportunity to make progress.

### 14.3 Packets and Virtual Channels (Decided)

Transactions are split into at most sixteen packets, each with up to
64 payload bytes. With a 16-byte header, the header accounts for
20 percent of transmitted bytes in full data packets.

| Virtual channel | Contents |
|---|---|
| request, real-time | classes 0 and 1 |
| request, normal | classes 2 and 3 |
| response, real-time | classes 0 and 1 |
| response, normal | classes 2 and 3 |

Real-time traffic may interrupt a normal packet at beat boundaries;
packet assembly and credits remain separate for each VC.
Four flit slots of eight bytes per VC yield 128 bytes of
link buffering per direction, in addition to header/assembler/tag state.
This is not a total area calculation for a router port.

In the 64-core target system, the crossbar does not use XY mesh routing.
XY routing belongs exclusively to future systems (section 23).
Protocol deadlock freedom requires separate credits, consumable
control responses, and the admission rules in section 27;
four VC names alone do not prove it.

### 14.4 Real-Time Classes

Intended order:

| Class | Use |
|---:|---|
| 0 | audio, hard real-time, error reports |
| 1 | display scanout |
| 2 | RAM DMA and normal task data |
| 3 | disk, USB, network, and background traffic |

Audio and display receive guaranteed time slots. Free time slots are
used by lower classes.

## 15. Shared RAM

### 15.1 Capacity (Decided)

The target system has 256 MiB of shared RAM, or 4 MiB per core, supplied by
exactly one DRAM component (section 15.2).

A useful measure of size is how quickly RAM can be filled
in practice: a USB-2 flash drive supplies about 35 MB/s and fills
256 MiB in about 7 seconds.

The 4 MiB per core is a capacity figure. It does not create fixed
private partitions for each core.

### 15.2 DRAM Component (Decided)

LPDDR4 or LPDDR4X is decided as a single-channel component with
2 Gbit x16 in a single die (SDP). The reference component is a
Winbond `W66BP6NB` in the x16 VFBGA100 version; an ISSI
`IS43LQ16128A` or its LPDDR4X variant is a second source to be
qualified.

```text
Organization:     1 channel x 16 DQ x 8 banks, 2 Gbit = 256 MiB
Package:          JEDEC-100-Ball-BGA
Target speed grade: 3200 MT/s, i.e. 6.4 GB/s theoretical peak
Operation:        3200 MT/s; for 3.2 GB/s full-data contact payload,
                  at least 50 percent DRAM efficiency is required
Reference:        Winbond W66BP6NB, x16, VFBGA100, 3200 MT/s
Second source:    ISSI IS43LQ16128A/AL, x16, BGA100, 3200 MT/s
```

Rationale:

- 2 Gbit gives exactly the target size of 256 MiB. Smaller densities exist
  from individual manufacturers but do not meet the decided capacity.
- One channel with 16 bits requires only about 34 to 38 signal pins, fewer than
  DDR3 or DDR4 x16.
- The theoretical DRAM peak exceeds the contact maximum.
  Sufficient sustained bandwidth for scanout and computation is a
  model/controller condition, not evidence already established.
- Winbond and ISSI offer suitable LPDDR4 families with long-term or
  industrial variants. The specific speed grade and temperature range must
  be checked against the orderable part numbers available at the time before
  layout approval.
- DDR3 was not selected as the target interface because of its higher pin count,
  higher I/O power, and long-term procurement risk.
- PSRAM (HyperRAM/OctalRAM) was rejected because configurations available with
  suitable capacity and pin count do not achieve the intended payload bandwidth.

LPDDR4 requires 1.8 V and 1.1 V supplies (LPDDR4X: VDDQ 0.6 V) and
a memory-controller training phase at startup. The DRAM PHY is
licensed IP (section 21).

Price, package option, temperature class, and availability must be
checked again before manufacturing. The architecture depends only on
2 Gbit, x16, and at least 3200 MT/s, not on a single
orderable part number.

### 15.3 Full Access

All cores must be able to address the entire shared RAM. Physical bank or
channel organization must not partition the visible address space.

The compiler and dispatcher allocate dynamic RAM resources with base,
length, and access rights. Within an assigned resource, a core uses
a 32-bit offset. The contact can nevertheless transport physical addresses
of at least 40 bits.

### 15.4 CPUlet-SRAM (Future Option)

T1 reserves an address window but has no additional
CPUlet-SRAM in the standard profile: size 0 bytes.
Tables, reference images, and shared structures reside in DRAM.
The cores' local SRAM remains unchanged.

256/512 KiB may be investigated in a later profile revision.
Even then, the same contact, explicit transfers, and
ownership rules apply; no cache is created.
A fixed end-to-end latency must not be inferred from an SRAM access time:
the router and competing requests remain relevant.
Area and benefit are demonstrated only in the corresponding profile.

## 16. Graphics Output

### 16.1 Decided Maximum

```text
Maximum output:               1280 x 720 at 60 Hz
Maximum internal color depth: 32 bits per pixel
```

1080p is intended for future multi-CPUlet systems (section 23). 1440p
and 4K are not targets.

### 16.2 Framebuffer

The framebuffer must be a normal region of shared RAM. There is
no separate VRAM and no parallel graphics memory path.

| Buffer at 720p | Size |
|---|---:|
| 32 bits, one image | 3.5 MiB |
| 32 bits, double buffer | 7.0 MiB |
| 32 bits, triple buffer | 10.5 MiB |
| RGB565, double buffer | 3.5 MiB |

| Metric at 720p, 32 bits, and 60 Hz | Value |
|---|---:|
| Scanout | 221 MB/s |
| 1-KiB transfers per image | 3600 |
| Link share for scanout alone, against 3.2 GB/s payload maximum | 6.9 % |
| Link share for scanout plus rewriting, same bottleneck direction | 13.8 % |

The sum is a link share only if the bottleneck direction is shared:
read/write data may physically use different directions.
At DRAM, both are data traffic; headers and pauses are counted separately.

A 720p triple buffer occupies about 4 percent of the 256 MiB.

### 16.3 Display Unit

The display unit is not a GPU replacement. It handles only
deterministic transport:

- framebuffer start address,
- line length,
- resolution,
- pixel format,
- switching image buffers during vertical blanking,
- request FIFO and line buffers,
- TMDS encoding and serialization for DVI-D.

The display unit is designed for single-link DVI up to a 165 MHz pixel clock.
720p is the deliberately decided product and quality limit of the target system.
The PHY could technically transport 1080p60 as well; however, the
target system still lacks measured end-to-end budgets for renderer, RAM,
contact, display, and simultaneous media load.

Tiles, sprites, text, composition, scaling, and effects remain core tasks.
T1 has no hardware scaler. An application scales explicitly through guest code
into the completed framebuffer; costs and filter selection
belong to its budget.

### 16.4 Digital Output: DVI-D

Single-link DVI-D (TMDS) is technically decided. Approval for a
production product remains subject to a final legal and
compliance review:

```text
3 differential data pairs
1 differential clock pair
= 8 signal pins
plus DDC (2 pins, I2C for EDID) and hotplug (1 pin)
```

| Mode | Pixel clock | Bitrate per data pair |
|---|---:|---:|
| 720p60 | 74.25 MHz | 742.5 Mbit/s |
| 1080p60 (future) | 148.5 MHz | 1.485 Gbit/s |

Rationale:

- No external protocol converter is intended. TMDS encoding is digital
  logic; however, serializer, line driver, ESD protection, and controlled
  output impedance form a characterized high-speed
  I/O macro and are part of the PHY and signal-integrity design.
- Interoperability with HDMI inputs via a passive DVI-HDMI adapter
  is a development goal, not a guarantee for every display device. EDID is
  read; unsupported modes are not output.
- Publication of the DVI-1.0 specification itself grants no
  IP license, but refers to a reciprocal royalty-free Adopter
  Agreement. Patent status, applicability of this agreement, and trademark use
  must be legally clarified before production approval.
- The device uses a DVI connector or a DVI-HDMI cable, not an
  HDMI connector. The HDMI specification, HDMI trademark, and HDMI logo remain
  excluded. DisplayPort remains excluded because of membership,
  document, and compliance dependencies.
- Audio is not transmitted over DVI; it has separate pins
  (section 17).

VGA was rejected: about 20 pins for 18-bit color through resistor ladders,
and modern monitors require active adapters. LVDS/OpenLDI was rejected
because it requires a dedicated receiver at the monitor.

1440p60 (about 241 MHz) would require dual-link DVI and does not work
through passive HDMI adapters. It is at most an option for
dual-link DVI monitors.

### 16.5 Rendering Model for 3D (Target Proposal)

The target is 3D in the class of Tomb Raider 1 at 720p, not modern 3D environments.

Estimate for 720p60: 1280 x 720 pixels with overdraw 2 yield
110.592 million pixels per second. An assumed 10 to 15 instructions per
textured pixel yield 1.106 to 1.659 GInstr./s, mathematically
about 5 to 7 fully utilized RUN cores. Six to eight with reserve
remain a hypothesis. Geometry, clipping, tile lists, texture misses,
and game logic must be budgeted additionally.

The bottleneck is texture access, not instructions. An `RREAD` per pixel
is ruled out. The renderer therefore works in tiles. Example
slot allocation for a core:

| Slots | Contents |
|---:|---|
| 4 KiB | image tile 32 x 32 at 32 bits |
| 2 KiB | tile Z buffer or polygon sorting |
| 2 KiB | current texture window 32 x 64 with an 8-bit palette |
| 3 KiB | code |
| 1 KiB | stack, descriptors, palette, and small state |

- Completed tiles are written to the framebuffer using `RWRITE`.
- The texture working set resides in DRAM; windows are loaded locally.
  The example allocation has no free double-buffer slot. Transfer and
  computation must therefore occur in phases, or the tile must be reduced.
- Lane modes (4 x 8 bits) serve Gouraud shading and blending.
- Perspective correction uses `CLZ`, a reciprocal table, and a Newton step.

## 17. Audio

Audio operates as a hard real-time stream. Up to 8 channels
(7.1) are intended.

Digital chip output without an external audio DAC (decided):

- PDM: one sigma-delta bitstream per channel. Interpolation (CIC),
  sigma-delta modulator, and the temporally characterized output cell reside
  on the chip.
- 7.1 requires 8 pins.
- In software, the modulator at about 3 MHz per channel would occupy nearly
  an entire core.
- The cores supply completed PCM samples through DMA.
- A passive RC low-pass filter provides only an analog experimental signal
  suitable for a high-impedance load. It is not a guaranteed line output and can
  drive neither headphones nor speakers. An analog output capable of driving
  its load requires at least a low-pass filter, buffer or amplifier,
  and protection circuitry outside the chip.
- Achievable quality is determined by modulator order, oversampling,
  output cell, I/O supply, clock jitter, filter, and load. The
  audio pins receive a separate filtered I/O supply; dynamic range,
  noise, and distortion are measured on the prototype rather than
  promised in advance as a bit count.

Alternative mode of the same audio unit on the same pins: I2S/TDM (TDM8 on
one data line) for boards with an external hi-fi DAC.

S/PDIF was rejected because it carries multichannel audio only in compressed,
license-requiring formats.

Audio bandwidth is small compared with graphics. Eight channels at 24 bits
and 192 kHz require about 4.6 MB/s; with 32-bit transport, about 6.1 MB/s. Audio
nevertheless receives the highest traffic class because latency and interruptions
matter more than absolute data volume.

One IDLE core for RAU or RFXL is a previous planning
hypothesis. Rate, image size, deadline, slot/RAM requirements, and
guest instruction costs must be demonstrated according to section 29.

## 18. USB and Mass Storage

The target platform uses USB for input devices, MIDI, and mass storage.
SATA, PCIe, USB 3, USB4, and Thunderbolt are omitted.

There is exactly one USB host logic with multiple ports and two PHY classes.

### 18.1 USB Low-/Full-Speed with Integrated PHY (Decided)

- Low- and Full-Speed (1.5 and 12 Mbit/s) use two data pins per port and
  an integrated, characterized LS/FS transceiver cell. It ensures, among
  other things, levels, receiver thresholds, output impedance, edge shape, and
  SE0 detection. Passive components and ESD protection required by
  the USB design are added on the board.
- Hardware: bit level, i.e. NRZI, bit stuffing, CRC5/CRC16, 1-ms SOF timing, and
  packet buffers.
- Software on the system cores: transaction and device level (HID, MIDI).
- One port per device. No hub chip is required.
- Target system: 3 ports for mouse, keyboard, and MIDI.
- The 5-V VBUS supply, current limiting, and overcurrent detection reside in
  external load switches or protection components and are controlled and
  queried by the chip.

### 18.2 USB 2.0 High-Speed (Decided)

- One High-Speed port (480 Mbit/s) for mass storage, typically a
  USB flash drive.
- High-Speed requires an analog PHY. It is integrated as licensed IP to
  avoid an additional chip. A USB-2 PHY is mature and substantially
  simpler than a USB-3 PHY.
- The same port also supports Full-Speed.
- Actual flash-drive throughput: about 35 MB/s.
- A USB-2 flash drive can operate on an LS/FS port in Full-Speed mode. Its
  usable throughput there is typically well below 12 Mbit/s; this
  suffices for many compressed media streams, but not for quickly loading
  large programs.

USB must block neither display scanout nor audio. Bulk transfers for disk
and network run in the lowest contact class.

### 18.3 Real-Time Decoding of Compressed Data

Compressed data is decoded directly during reading:

```text
Drive --USB--> RAM ring buffer --Contact cl. 2--> Decoder core slot --> Audio / Framebuffer / RAM
       (cl. 3)                    (double buffer: slot A fills, slot B decodes)
```

- A configured RAM ring buffer of, for example, 1 to 2 seconds
  bridges bounded storage pauses. It provides no guarantee against
  unbounded USB stalls. The profile specifies bitrate, buffer bytes, and
  maximum permitted pause.
- Parallelization follows the existing bitstream boundaries and
  dependencies of each format family (section 29). RAU frame copies,
  RFXL predictors, and PMF0 history are not treated as independent.
  Decoder and USB can both be bottlenecks.
- Example: 35 MB/s compressed at a factor of 3 yields over 100 MB/s of payload.

## 19. Error Handling, Boot, and Debug (Decided)

There are no interrupts or exceptions in the classical sense.

Fault causes:

- invalid opcode,
- access outside an assigned resource,
- slot ownership conflict.

Sequence:

1. The core enters the `FAULT` state; its state is preserved.
2. It sends a class 0 fault message to a designated
   system core. The message value contains cause and flags.
3. The system core reads PC, registers, and status through status registers
   addressable through the contact.

The same mechanism serves boot and debug: a system core can halt, inspect,
load, and start any core through the contact. There is
no parallel debug path.

## 20. Package and Pins

Preliminary minimum estimate for 64 cores, one LPDDR4 component,
DVI-D, PDM audio, and USB:

| Area | Pins or balls |
|---|---:|
| LPDDR4, one channel x16 | about 34 to 38 |
| USB 2.0 High-Speed including port control | about 4 to 5 |
| USB Low-/Full-Speed, 3 ports including VBUS control | about 9 to 12 |
| DVI-D including DDC and hotplug | 11 |
| Audio PDM or I2S/TDM | 8 |
| Clock and reset | about 3 |
| Boot and minimal debug | about 2 to 4 |
| Calibration and references | about 2 to 4 |
| **Functional signals** | **about 73 to 85** |
| Power and ground | about 40 to 55 |
| **Total** | **about 113 to 140** |

A 144-ball BGA remains the preferred lower bound, but this functional
calculation leaves only about 4 to 31 spare balls. Whether power,
return-current paths, LPDDR4 escape routing, and separation of sensitive I/O
domains can be accommodated reliably must be shown by a specific package and
pinout study. If the reserve is insufficient, 169 BGA is the next
target size; this does not change the architecture itself. The core signals
do not leave the chip.

## 21. Rough Area and Energy Guidance

Without a target process, SRAM macros, synthesis, and physical layout, no
reliable area or power figures are possible.

An earlier rough 28-nm estimate for a core still envisaged as capable of
750 MHz yielded:

- about 35 to 60 kGE of logic,
- about 0.06 to 0.12 mm2 including 12 KiB SRAM,
- about 3 to 8 mW at 250 MHz as a rough activity scale.

The final design limited to 250 MHz should become smaller and more energy-efficient
because the fast clock tree, 1.33-ns critical paths, and 750-MHz voltage level
are eliminated. A CPUlet-SRAM with 256 KiB would add roughly 0.4 mm2. These
values must not be used as specifications until actual PPA synthesis
has been performed.

Regarding the manufacturing process: open PDKs (SKY130, GF180MCU, IHP SG13G2)
fit the open design ambition, but barely reach 250 MHz with single-cycle
multiplication, and for 64 cores yield roughly 80 to 170 mm2 and do not offer
the required LPDDR4, USB-2-HS, USB-LS/FS, and TMDS I/O macros as
ready-made, characterized standard solutions. A commercial process
around 28 nm with licensed or specifically qualified PHY IP is realistic. ISA,
RTL, and protocols can remain open; manufacturing, PHY macros, and the use
of external standards must be assessed separately.

## 22. Future Physical Prototypes (Outside the T1 Lab)

An FPGA is a possible future investigation before an ASIC.
The current project scope remains complete theory followed by
a Mac lab app; an FPGA is neither purchased nor required for it.
The following historical guidance figures are not a promise.

| FPGA | Toolchain | Cores (estimate) | Clock (estimate) |
|---|---|---:|---:|
| Lattice ECP5-85 | fully open | about 16 to 24 | about 80 to 120 MHz |
| AMD Artix-7 200T | Vivado, free of charge | about 32 to 48 | about 100 to 150 MHz |
| Kintex class | Vivado | 64 | - |

The estimate is based on 2 to 4 k LUTs per core and must be confirmed by the
first synthesis. A partial configuration with 16 to 32 cores
is sufficient because the ISA is independent of core count.

Rules:

- The contract applies in cycles: one instruction per cycle and fixed branch times.
  With proportional scaling, RUN:Fabric=1:2 and IDLE:Fabric=1:10 remain,
  for example 20/100 MHz for IDLE/RUN and 200 MHz fabric.
  Instruction budgets then remain comparable; actual RAM/PHY latencies,
  image/sample rates, and absolute real-time deadlines must be budgeted anew.
- FPGA boards usually have DDR3, and ECP5 and Artix-7 do not directly support
  LPDDR4. The prototype may therefore use DDR3. The memory controller
  sits behind the contact; the contact contract is identical.
- DVI, PDM audio, and USB LS/FS can be functionally prototyped with suitable
  FPGA I/O cells. This does not yet demonstrate USB compliance or
  TMDS signal integrity; if the board I/Os do not meet the electrical
  requirements, an external transceiver solution is used for the prototype.
  USB 2.0 High-Speed requires an external PHY chip with a ULPI connection
  on the FPGA.
- The emulator is first brought into line with the target behavior in this
  document and remains the sole reference model. RTL is checked against it
  in co-simulation.
- The FPGA is used to measure whether CPUlet-SRAM is needed and, if so,
  at what size.

## 23. Future: Multi-CPUlet Systems (Open)

This section is not part of the target system. It collects previous
considerations so that current decisions do not prevent future larger systems.
None of this is decided.

Guidelines for the target system:

- The ISA remains independent of core count.
- The sender identifier of messages (16 bits) leaves room for a
  CPUlet identifier.
- The contact transports addresses of at least 40 bits.
- The contact contract also applies unchanged between CPUlets.

### 23.1 Macroblock and Maximum Configuration

A macroblock arranges 16 CPUlets as a 4-by-4 mesh (1024 cores). Each
CPUlet receives four additional mesh ports (north, south, east, west);
the crossbar grows from 10x10 to 14x14. Simple deterministic
XY routing is preferred. Four macroblocks (2 x 2) yield the previous
upper limit of 4096 cores.

| Cores | CPUlets | Macroblocks | RAM at 4 MiB per core | Total local SRAM | Display output |
|---:|---:|---:|---:|---:|---|
| 64 (target system) | 1 | - | 256 MiB | 768 KiB | 720p60 |
| 256 | 4 | - | 1 GiB | 3 MiB | 1080p60 |
| 1024 | 16 | 1 | 4 GiB | 12 MiB | 1080p60 |
| 4096 | 64 | 4 | 16 GiB | 48 MiB | 1080p60 |

Theoretical peak performance at 250 MHz: 64 GInstr./s with 256 cores,
256 GInstr./s with 1024, and 1024 GInstr./s with 4096 cores.

### 23.2 RAM and Banking

A single RAM port is insufficient for large systems. Previous
consideration for a 1024-core macroblock:

- 16 logical RAM banks,
- multiple physical memory controllers,
- interleaving at 1-KiB slot boundaries,
- every address reachable from every CPUlet.

As an initial working assumption, four physical memory controllers with four
logical banks each were considered. This means 256 rather than 64 cores share
a controller; bandwidth per core drops to one quarter of that in the
target system. In addition to capacity, a bandwidth rule per core
would therefore need to be defined.

### 23.3 Input and Output

- From 4 CPUlets onward, 1080p60 is intended through the same single-link DVI
  display unit. A 1080p framebuffer at 32 bits: 7.9 MiB per image, scanout
  498 MB/s, 8100 1-KiB transfers per image, 15.6 percent of the
  3.2-GB/s full-data payload maximum without additional load.
- Larger variants require larger packages because of more memory channels
  and higher current.

## 24. Differences Between the v4.1 Emulator and Hardware Target

The emulator is a valuable functional starting point, but not a
cycle-accurate hardware model.

| Topic | v4.1 emulator | Target hardware |
|---|---|---|
| Core count | models 4096 | 64, one CPUlet; multi-CPUlet systems are future work |
| Cluster | 64 cores | 64 cores equal one CPUlet, 8 groups of 8 |
| Clocks | symbolic `ULTRA/MIDI/TURBO` | 0/50/250 MHz, no 750 MHz |
| Execution | one call per simulation quantum | one instruction per active cycle |
| Pipeline | not cycle-accurate | four stages; untaken branch 0, taken branch 2 additional issue cycles |
| Local memory | 12 KiB, lazily allocated | 12 x 1 KiB banks with slot ownership |
| External RAM | not a core component | shared RAM, 256 MiB |
| Load/store | separate 16/32-bit opcodes | unified 8/16/32 bits |
| Constants | `MOVI` with 14 bits | `MOVI` with 22 bits plus `MOVHI` |
| Resources | block and sleep | nonblocking, response in local slot, completion as a message |
| Messages | FIFO without delivery acknowledgment | Tx buffer, ACK/NACK, automatic retry |
| Errors | not modeled | `FAULT` plus fault message |
| Graphics | lab/audit environment | DVI-D display unit with RAM framebuffer, 720p60 |
| Audio | model resource | PDM or I2S/TDM, up to 7.1 |
| Peripherals | model resources | uniform contact, integrated USB-LS/FS and USB-2-HS PHYs |

New hardware decisions should therefore be reflected first in this specification
and then specifically in the emulator. Old behavior must not persist as a
second permanent architectural path.

## 25. Decided Theory and Open Physical Validation

ISA, compiler concept, ABI, resource/message contract, format profiles,
and lab contract are decided for T1 in sections 9..14 and 27..31.
The theoretical specification no longer leaves an optional LOOP, unknown
signed MUL, or unspecified transfer format open.

The following remain open for a future physical implementation:

1. Process, cells, SRAM/128-bit descriptor banks, and PHY/I/O macros.
2. Timing of MUL/MAC, SRAM, forwarding, and resource acceptance.
3. Clock domains, reset/wake transitions, and power supply.
4. DRAM controller, training, refresh, and sustained bandwidth.
5. Electrical boot source and package/pinout study.
6. USB/TMDS compliance and PDM output quality under load.
7. DVI legal/Adopter review before production approval.
8. Measured area, power, and temperature budgets.

Compiler, guest codecs, and reference model still need to be implemented or
aligned according to this theory. Their functional, timing, and worst-case
validation follows the contract in section 31;
this work has not been completed in this round.

## 26. Compact Target Summary

The uRISC target system is a 64-core system consisting of one CPUlet with eight
groups of eight cores. Four cores are reserved for system and I/O tasks,
60 for applications. Each core has twelve 32-bit registers,
four resource connections, a 64-bit MAC accumulator, a message FIFO,
and twelve local 1-KiB slots with exactly one owner per cycle. Cores
operate at 0, 50, or at most 250 MHz.

The four-stage pipeline executes one instruction per cycle. Normal instructions
are deterministic and do not wait for each other; taken branches cost a fixed
two additional issue cycles. External work is organized through nonblocking
requests with reserved response space. Responses land in local SRAM;
completions, events, and errors arrive as messages, for which only an
explicit `WAIT` waits. The compiler splits programs into tasks, slots,
and data regions during compilation itself.

Shared RAM comprises 256 MiB, or 4 MiB per core, from a single
LPDDR4 component with 2 Gbit x16. All cores can address the entire RAM.
The framebuffer resides exclusively in this RAM. CPUlet-SRAM is
reserved as an address window but in the T1 standard profile is 0 bytes.

A uniform 64-bit contact at 500 MHz, with transactions up to 1024 bytes,
64-byte packets, and four virtual channels, connects core groups, RAM,
DMA, display, audio, and USB. Audio and display receive guaranteed real-time windows.

The target system outputs at most 720p60 through DVI-D. Interoperability with
HDMI inputs through a passive adapter is intended and verified using
EDID and prototypes; production approval requires legal and electrical DVI
review. Audio runs at up to 7.1 as PDM without an external DAC or alternatively
as I2S/TDM; an analog output capable of driving its load
requires external filter and output stages. Mouse, keyboard, and MIDI connect
to three ports with integrated USB-LS/FS PHYs, mass storage to one port with
an integrated USB-2.0-High-Speed PHY. Proprietary formats execute according to
their existing stream/packet/root boundaries and dependencies. HDMI,
DisplayPort, SATA, PCIe, USB 3, USB4, and Thunderbolt are not part of the
target architecture.

Systems with multiple CPUlets up to 4096 cores remain open and future work
(section 23). They should use the same ISA, cores, and
contact contract.

## 27. System, Memory, and Event Contract (Normative)

### 27.1 Configuration, Numbering, and Limits

The T1 standard profile has 64 cores and 256 MiB RAM. The visible
8x8 arrangement numbers row y and column x as `coreID=8*y+x`.
Each fabric group is a row of eight cores.
The screen position is a representation, not an eight-way meshed
physical connection. Cores 0..3 are system cores.

| Quantity | Decided value |
|---|---:|
| local bytes per core | 12288 |
| GPRs | 12 x 32 bits |
| ACC | 64 bits |
| resource bindings | 4 |
| local slot count | 12 |
| Rx FIFO | 16 x 64 bits |
| of which normal input slots | 8 |
| of which completion slots | 8 |
| Tx buffer | 8 entries |
| simultaneously open RREAD/RWRITE | 8 |
| transaction identifier on contact | 16 bits |
| shared RAM size | 268435456 bytes |
| CPUlet-SRAM in standard profile | 0 bytes |
| system tick | 500000000 ticks/s |
| RUN / IDLE clock | 250 / 50 MHz |

An Rx FIFO contains the combined accepted sequence of both classes.
The quotas are admission limits, not two separately readable FIFOs.
A resource request reserves one of the eight completion slots.
The slot remains occupied until RECV consumes the completion, even if the
transaction has already completed. There are therefore never more than
eight unconsumed resource completions.

### 27.2 Visible States and Reset

The core has:
OFF, READY, RUNNING, WAITING, STOPPED, and FAULT.
OFF is power-/clock-gated and cannot be awakened by normal messages.
READY is prepared but not yet started. RUNNING operates at
the configured frequency; IDLE is a frequency, not a state.
WAITING retains registers, ACC, PC, bound resources, and all
required slots. A message or timer can wake it.
STOPPED and FAULT require an explicit dispatcher action.

Reset sets GPRs, ACC, PC, FIFO/Tx/tag state, and timers to zero,
revokes all resource rights, and sets cores 1..63 to OFF.
Core 0 starts only after a validated boot task image is loaded.
SRAM and DRAM may be physically uninitialized; the loader
initializes every readable region before enabling it.
The model treats access to uninitialized regions as
FAULT_UNINIT. It supplies no host-dependent random bytes.

The theoretical boot path is an externally supplied boot image:
loader -> local code of core 0 -> RAM initialization ->
system tasks 1..3 -> application tasks. The eventual electrical boot source
has no influence on the theoretical specification; the app will
load the same validated image across the host boundary.

### 27.3 Slot Rights, Owners, and Code Changes

Each slot has the role CODE or DATA and an initialization state.
A role is set only by the dispatcher while the core is halted.
DATA may additionally be DMA-owned during a transfer.
The owner per cycle is IF, EX, or DMA, never more than one.

Persistence means retaining contents, not granting access to others.
A WAIT with open transfers must lose neither descriptor/data state nor
Rx/tag state. A transition to OFF requires that no
transfers or message acknowledgments remain open.
Control logic for wake and reception remains reachable in WAITING.

Self-modifying code is not allowed in a running task.
Code overlay:
task halts -> dispatcher waits for open transfers ->
affected bank becomes DATA -> RREAD and completion ->
code validation -> CODE role -> start at an enabled PC.
No asynchronous code change during IF.

The sum of code, data, stack, and DMA banks must fit in twelve slots.
A 12-KiB working set with no room for code/stack is not a valid task image.
The CPUlet's total SRAM capacity is 768 KiB.

### 27.4 Resources and Access Rights

S0..S3 are opaque bindings to validated resource entries. A
task cannot change their bits with MOV. The dispatcher binds
base address, length, READ/WRITE rights, generation, owner, and
traffic class. The physical base has up to 40 bits, the descriptor
a 32-bit offset. A T1 resource window is at most 256 MiB.

An offset/length pair is checked with `offset <= size` and
`length <= size-offset`, never through an overflowing addition.
A relinquished or reassigned resource receives a new generation.
Generation exhaustion retires the ID instead of wrapping and revalidating an
old token. The same rule applies to task, endpoint and destination generations.
Old bindings return CAPABILITY and must not trigger RAM access.

Resources are RAM windows or memory windows of a device service.
Control commands for display/audio/USB are written as validated descriptors
into such windows. A core has no second unrestricted
MMIO access and cannot arbitrarily address other cores
or device registers.

System resources for debug, clock, and task control are bound exclusively to
system cores or the halted host loader.
Application code receives only the required data/service windows.
A classical MMU is not required: LD/ST remain local,
external accesses are exclusively capability-checked.

### 27.5 Status and Messages

Instruction status is a u32:

| Value | Name | Meaning |
|---:|---|---|
| 0 | OK | accepted or successfully read |
| 1 | BUSY | Tx, tag, completion quota, or required data slot occupied |
| 2 | EMPTY | RECV without a message |
| 3 | TARGET | destination identifier invalid or destination permanently unreachable |
| 4 | DESCRIPTOR | length, alignment, or local descriptor region invalid |
| 5 | CAPABILITY | resource unbound, stale, or lacking rights |
| 6 | RANGE | resource range invalid |
| 7 | IO | accepted transfer fails at the device |
| 8 | CANCELLED | accepted transfer terminated by system cancellation |

Normal rejections change only the status register and nextPC.
An accepted transfer reports a later IO/CANCELLED completion
instead of a second local return value. After a partial error,
the entire RREAD destination must be treated as invalid; no partial
success is assumed. RWRITE is not atomic in external RAM.

The 64-bit message consists of source16, kind8, status8, value32.
USER=0, READ_DONE=1, WRITE_DONE=2, FAULT_EVENT=3, SYSTEM=4, TASK_DONE=5.
USER has status zero; value is the SEND value. READ_DONE/WRITE_DONE
carry the descriptor cookie as value and the completion status.
FAULT_EVENT has source=coreID, status=0, and value=fault code;
the privileged debug service reads fault PC and instruction.
TASK_DONE has source=coreID, status=0, and value=task ID.
Core sources 0..63 are USER/FAULT_EVENT/TASK_DONE; resource sources FF00..FF03
denote the bound S0..S3, FFFE the dispatcher.
The remaining identifiers are reserved for future systems.

Cookies are unique among a task's open requests.
The tag is reused only when transfer and completion consumption
have ended. A data slot is released before its
completion message becomes visible. The core may read or transfer it again
only after RECV of the matching successful completion.

ACK/NACK are fabric protocol, not USER messages.
An ACK confirms insertion, not RECV. Tx sequence identifiers are
not reused while outstanding. A destination contains
a finite dedup table for the outstanding windows of all sources;
the specific memory organization is an implementation matter.
Source-destination order is maintained even with NACK and retry.
T1 models internal links as lossless: an accepted packet
does not disappear, ACKs are not lost. A diagnosed
link/model error ends the experiment with DEVICE/INTERNAL.
Retry is triggered only after NACK, not by an invented
ACK timeout. Sequences run modulo 65536 per source-destination pair;
at most eight are outstanding, and the destination retains the last
eight accepted sequences as its dedup window. Old in-flight packets
must not survive a destination generation change.
A later model with packet loss requires its own
protocol revision and must not silently replace this assumption.

A faulted or stopped receiver is not blindly sent data forever:
system cancellation invalidates the destination generation and terminates affected
Tx entries with a diagnosable error. Such an error
causes FAULT_DELIVERY at the sending core because SEND has no reserved
register slot for a later return value. Fault state is latched separately
and cannot be lost because of a full Rx FIFO.

### 27.6 Ordering, Visibility, and Completion

There is no general implicit RAM fence. The contract is:

1. A successful RWRITE completion means all bytes are visible to
   subsequent accesses by other masters.
2. Only then does a message transfer data ownership.
3. The receiver consumes the message and may start RREAD.
4. A successful RREAD completion means all destination bytes have been
   written locally and are visible.

Requests from a core may complete out of order if there is no dependency;
cookies identify them. Two simultaneously writing
tasks with overlapping regions are an invalid ownership plan.
An RWRITE that is only buffered in the controller must not yet
trigger a visibility ACK.

Acceptance, linear byte access, and global transaction atomicity are
different. A 1024-byte transfer is not atomic; data ownership
prevents observation of partly written regions. Atomic RAM ALU instructions
are not part of T1.

### 27.7 Contact Packet and Actual Payload Bandwidth

The raw bandwidth of 64 bits x 500 MHz is 4 GB/s per direction.
For T1, a packet header of 16 bytes is decided:

| Bit range of the 128-bit header | Field |
|---|---|
| 39:0 | physical address or service offset |
| 55:40 | source |
| 71:56 | destination |
| 87:72 | tag/sequence |
| 94:88 | payload length 0..64 bytes |
| 98:95 | kind |
| 100:99 | traffic class 0..3 |
| 105:101 | status |
| 109:106 | packet index 0..15 |
| 113:110 | packet count minus 1 |
| 121:114 | byte mask |
| 127:122 | flags, in T1 only bit 122=LAST allowed |

Kind: 0 READ_REQ, 1 READ_DATA, 2 WRITE_DATA, 3 WRITE_ACK,
4 MESSAGE, 5 ACK, 6 NACK, 7 ERROR; 8..15 reserved.
The byte mask applies only to a short single beat; for more than eight
payload bytes it is FF, and payload length limits the final beat.
Zero payload length applies only to pure control packets.
For READ_REQ, length denotes the requested packet payload;
no payload beats follow there. All other data types provide
exactly ceil(length/8) beats; unused final bytes are zero.

A full data block has 2 header beats + 8 payload beats = 10 beats = 20 ns.
A 1-KiB transfer has 16 data packets = 160 beats = 320 ns without
arbitration/response. Its full unidirectional bandwidth is therefore
3.2 GB/s rather than 4 GB/s; headers are 20 percent of total bytes.
READ requests and WRITE completions additionally occupy the opposite direction.
Small transfers, control traffic, and arbitration further reduce
payload bandwidth. Raw and payload values are never equated.

Real-time traffic may interrupt normal data only at beat boundaries. An
interrupted packet retains parser state, VC, and tag; the other
class has separate packet assembly. Packets of the same VC are
not interleaved with each other in the middle of a payload.
Request/response have separate credits. An endpoint must be able to
consume pure ACK/NACK/ERROR even under data pressure.

### 27.8 Arbitration and Real-Time

Group arbiters use round robin per class, routers also per output.
In 100 fabric ticks, the standard model reserves:
4 ticks for class 0, 12 for class 1, 64 for class 2,
20 for class 3. Unused ticks are available to other classes;
minimum quotas remain under competing load.
Audio/display data uses packets of at most 64 bytes.
Real-time-class control traffic is rate-limited so that faulty
software cannot claim the media quota without bounds.

At a continuously usable output, the reservation theoretically provides
up to 128 MB/s for class 0 and 384 MB/s for class 1 with full
data packets. Other traffic in the same class and
controller pauses must be subtracted. Audio requires only a small fraction;
720p60 display requires 221.184 MB/s of active pixel data.
A link quota is not yet a DRAM worst-case proof.

RAM models refresh, bank conflicts, and access latencies as
explicit events. Hard real-time is proven only for a profile
that bounds worst-case pauses and permitted traffic loads.
Mass storage with unbounded delays has no hard
real-time guarantee; buffers reduce failure risk but do not eliminate it.

### 27.9 Errors and Cancellation

Precise fault codes:

| Value | Code | Cause |
|---:|---|---|
| 1 | ILLEGAL | invalid opcode, field, or mode |
| 2 | PC | invalid fetch/branch target |
| 3 | ALIGN | misaligned local access |
| 4 | LOCAL_RANGE | local access outside SRAM |
| 5 | SLOT | role/ownership conflict |
| 6 | UNINIT | uninitialized readable region |
| 7 | INFLIGHT | DONE/OFF with open operations |
| 8 | DELIVERY | accepted message undeliverable |
| 9 | DEVICE | fatal device error not treatable as transfer status |
| 10 | INTERNAL | violation of a model invariant |

Fault PC and instruction word are latched. Older instructions
remain completed; the faulting instruction and all younger ones have
no local side effects. External errors after acceptance are reported as
transfer status or later DELIVERY, not retroactively projected onto
an old PC.

A running DMA retains its slot on fault and completes the transfer
or produces CANCELLED. Debug can read the slot only afterwards.
External writes already performed are not rolled back.
Fault notification to the system service is retried from a reserved
error latch until accepted; status inspection remains possible.

Cooperative cancellation stops new requests, drains completions/Tx,
returns ownership, and ends the task. Forced cancellation is saved in the
Scrapbook as cancelled, not as a successful experiment.

### 27.10 Device Services and Pixel/Audio Contract

Display presents only a fully written buffer.
PRESENT supplies resource binding, offset, width, height, stride,
format, and sequence number. It switches at VBlank and reports the
actual presentation as a SYSTEM message.
Only one pending present entry is allowed; further ones return BUSY.
The old front buffer becomes writable again only after the switch.

RGBA8 denotes the byte sequence R,G,B,A. An LD32 therefore reads
0xAABBGGRR; internal notation must not suggest ARGB.
Scanout ignores A after prior composition. RGB565 is
little-endian with R[15:11], G[10:5], B[4:0].
Stride is in bytes, at least width*pixelbytes, and 8-byte-aligned.
Framebuffer addresses are aligned to 64 bytes.

Audio takes interleaved PCM from a RAM ring:
1..8 channels, signed 16 or signed 24 in 32 bits, explicit sample rate,
channel order FL,FR,FC,LFE,SL,SR,BL,BR.
Mono/stereo use the front entries. 24 bits are right-aligned with correct
sign extension in the 32-bit word; PDM/I2S pack from this.
Sample rates 8000,16000,24000,32000,44100,48000,96000,192000 are
service profiles, not a guaranteed RAU codec limit.

The central audio mixer produces a final PCM stream, not
format-specific direct connections to the DAC. Underflow supplies silence,
counts an event, and causes a real-time check to fail.
Display underflow repeats the last valid image or
supplies a defined error state, never unchecked RAM.

USB supplies bytes and time events into RAM ring buffers.
Parsers receive bounded memory windows, not host file pointers.
The theoretical USB service models throughput and pauses; it
claims no electrically validated USB implementation.


### 27.11 Uniform Service Commands

The following structures are T1 runtime contracts, not new
codec file formats. A service window is bound only to an authorized
task. RWRITE at offset 0 supplies a command and
payload as a contiguous local memory block.

The 32-byte command header is little-endian:

| Offset | Type | Meaning |
|---:|---|---|
| 0 | u16 | protocol version, 1 |
| 2 | u16 | service operation |
| 4 | u32 | command cookie |
| 8 | u32 | payload bytes, 0..992 |
| 12 | u32 | flags, T1=0 |
| 16 | u32 | task ID of the authorized caller |
| 20 | u32 | resource/task generation |
| 24 | u32 | reserved, 0 |
| 28 | u32 | reserved, 0 |
| 32 | bytes | payload of exactly the declared length |

Header and payload remain within a 1-KiB slot. Service identity
follows from the binding, not from an arbitrarily selectable global device pointer.
Task ID/generation are checked against the bound caller identity.
An error after transfer acceptance produces an erroneous WRITE_DONE;
it must not execute a partly validated command.

WRITE_DONE=OK confirms insertion into the bounded service queue.
A longer-running service later generates SYSTEM with the
command cookie and status. Any result is visible beforehand in the
bound result window. The normally bounded Rx quota
uses ACK/NACK like other service messages; the service retains
pending events. Admission limits service queues to eight commands
per calling task. Full queues respond to the write with BUSY.
A task must distinguish both stages: WRITE_DONE and service result.

Payloads contain logical resource IDs, not physical pointers.
The service resolves ID plus generation according to the caller's rights.
IDs are allocated only by the dispatcher. The four S bindings suffice
for data traffic; a task must not expand its rights through a command.

| Service / Operation | Payload and result |
|---|---|
| System / 1 GET_TIME | no payload; result window: time u64, coreID u32, wakeReason u32 |
| System / 2 GET_TX_STATE | no payload; result: occupied Tx u32, open tags u32, unread completions u32, reserved u32 |
| System / 3 DIAGNOSTIC | up to 992 validated UTF-8 bytes; SYSTEM=OK after insertion into the bounded diagnostic ring |
| Dispatcher / 1 START_TASK | task image ID, generation, coreID, entry PC as four u32 |
| Dispatcher / 2 STOP_TASK | task ID, generation as two u32; orderly halt/cancellation |
| Display / 1 PRESENT | resource ID, generation, offset, width, height, stride, format, sequence as eight u32 |
| Audio / 1 CONFIGURE | resource ID, generation, offset, ring bytes, rate, channels, sample format, sequence as eight u32 |
| Audio / 2 SUBMIT | sequence, write position, valid frames, reserved as four u32 |
| USB / 1 READ_BLOCKS | resource ID, generation, offset, byte count, logical block index u64 as 24 bytes |

GET_TIME freezes the time at service acceptance. The result window
is located in the respective system service starting at offset 1024, has 16 bytes,
and is disjoint for each caller. It may be read using RREAD only after
SYSTEM=OK. Until that read completes, no second GET_TIME/
GET_TX_STATE is allowed on the same binding. Thus two LD instructions
cannot tear a 64-bit time, and no later command overwrites unread data.
wakeReason: 0=no wake since task start, 1=message, 2=timer,
3=both in the same tick; GET_TIME does not consume the reason.

Display format: 0=RGBA8, 1=RGB565. Audio format: 0=PCM16,
1=PCM24-in-32. CONFIGURE is possible only without an old running stream.
SUBMIT is monotonic modulo ring size and reports only fully
visible audio frames. The service must not release a ring region still being
read before the read cursor has passed it.
USB block index and sizes are checked against the modeled medium.

START_TASK/STOP_TASK are privileged; an application task reaches
them only through the designated system dispatcher. The header is a
validation boundary, not a security substitute for resource rights.

## 28. Compiler, ABI, and Runtime (Normative as a Concept)

### 28.1 One Frontend, One uRISC Target

Smallsome 3.0, in Code or Symbols source mode, is the language. The existing
[frontend](../../Ozon/module_cymple_front.pbi) is reused;
a future `urisc-t1` target is introduced alongside the existing Ozon target
without copying the grammar or parser.
Lowering, resource analysis, and code generation are target-specific.
The ISA may also be used directly through the assembler.

Both source modes normalize into one AST/typed IR. The current frontend
implements only its established Code subset; Symbols and the expanded 3.0
grammar require future work in that same frontend. The complete language's
binary64 num is not redefined by T1: the table below is the explicitly
restricted integer profile, not full-language conformance. Endpoint directions,
ownership and bounded admission follow the shared language specification.

A compiler project contains source text, target profile, input resources,
input/output contracts, and declared budgets. These produce
machine code, slot images, task graph, resource plan, and proof report.
An executable image without this metadata is not an approved
real-time program.

### 28.2 Supported Language

| Language feature | T1 contract |
|---|---|
| word | native signed 32 bits; addition/subtraction/multiplication modulo 2^32 |
| num | binary64 language semantics; only proven equivalent bounded integer operations, otherwise rejected |
| bool | canonical 0 or 1; CMP masks kept separate internally |
| fixed arrays, structs, bytes | static size, local or bound RAM storage |
| text | UTF-8 literals and existing LeXA resources |
| fn, if, match, for, while | normal control-flow lowering |
| on start / on frame | one-time initialization task or task bound to VBlank |
| print | bounded UTF-8 diagnostic block through system service, counted guest work |
| spawn, channel, send, receive | bounded tasks, fixed channel/message capacities |
| borrow/move | statically checked ownership across local and RAM regions |
| word / and % | shared software division, truncated toward zero; no DIV opcode |
| list/map, dynamic text | general dynamic values unsupported; bounded quantum outcome storage is statically reserved |
| extern | no guest-host function calls; resources instead of native libraries |
| race/collect/timeout | bounded task set, cancellation/join plan |
| recursion | only with statically bounded depth; otherwise rejected |
| floating point | unsupported in T1; explicit fixed-point libraries |
| detach, host file/network, arbitrary timer allocation | rejected by the static real-time profile |

Changing targets never changes num into word. num 1/2 cannot lower to integer
zero; a num addition whose exact binary64 result exceeds the declared lowering
range is rejected rather than wrapped. word arithmetic explicitly requests
the machine contract in either source mode. Fixed-point formats are documented
word-based library contracts (width, sign, fraction bits), not ungrammatical
new source types. Explicit to_word/to_num conversions preserve the language's
checked boundary. Nullable fields require initialized presence tags or proof
that null is unreachable; fixed arrays may not silently replace null by zero.
Range induction stops before exceeding its mathematical bound, not by wrapping
INT_MAX back to INT_MIN. A bounded-target report accounts for that check.

A language feature not yet supported by the frontend remains a
specific extension task for the same parser. This text does not claim
that the existing Ozon backend already has these T1 capabilities.

Runtime errors such as division by zero are reported as defined task errors
to the dispatcher. Hardware FAULT is a machine error and
is not automatically treated as a catchable Smallsome Code exception.
The exact Smallsome 3.0 language semantics are retained when aligning the backend;
if language and target arithmetic differ, lowering must
explicitly compensate for the difference or reject the program.

Collection indices remain 1-based according to Smallsome Code. Lowering checks the
source index before converting (index-1)*elementsize into a byte address.
struct field offsets and padding are recorded in the build report, not in the
bitstream of an embedded format.
A new on-frame event does not start a second overlapping writer
of the same frame state: if the task is still running, the profile reports a
deadline violation and skips the restart.
on-start completes before normal application tasks begin.

guru/throw/rethrow are implemented as explicit control-flow/task-error edges
according to Smallsome Code semantics. A handler requires a statically
bounded context and a known target; it is not a hardware IRQ.
An error in a handler that can no longer catch it ends the task with ERROR.
print serializes only permitted values into a bounded
diagnostic buffer; unbounded string allocation is not allowed for this.

### 28.3 Compilation Steps

1. The existing frontend produces the syntax tree and source positions.
2. Type checking resolves sizes, resources, and value ranges.
3. A typed intermediate representation models basic blocks,
   phi values, effects, and explicit ownership transitions.
4. Bounds, alias, and dependency analysis determines read and
   write intervals and necessary task edges.
5. Parallelization splits only proven independent work.
6. Slot planning assigns code, data, stack, and ping-pong transfers.
7. Register allocation and ISA lowering produce T1 instructions.
8. Scheduling defines cores, clock level, transfers, and budgets.
9. The linker checks branch targets, overlays, resources, and the overall image.
10. A report states assumptions, limits, worst-case paths, and open
    runtime conditions. Only then is the image considered admissible.

The intermediate representation is a compiler product, not a second
executable machine. Only T1 machine code is executed.

### 28.4 Parallelization and Granularity

A loop is split only if:
iteration bounds are bounded, write intervals are disjoint,
read data is stable, and no unhandled dependency exists.
Pointer aliasing, variable indices, and successive LZSS/predictor states
may prevent splitting. Execution then remains serial;
an assumed speedup does not count as proof.

Preferred task sizes are whole image tiles, PMF0 roots, bounded
audio sections, or existing resource chunks. A task per pixel or
sample is ruled out because of message and transfer overhead.

Reductions receive local partial results and an ordered join.
word modulo addition may be reordered after algebraic proof;
signed saturation is generally not associative.
The compiler must not freely reorder a saturating reduction.
Identical results must not depend on the incidental number of host threads.

### 28.5 Task Graph and Dispatcher

A task descriptor contains:
task ID, code image and entry PC, parameter/stack layout, slot roles,
four resource bindings, permitted cores, clock level, dependencies,
deadline, WCET budget, and cancellation path.

A task becomes READY only when all inputs are visible and its
destination core/resource plan is free. The dispatcher starts READY tasks.
Normal application tasks are run-to-completion; WAIT permits asynchronous
continuation of the same task, not automatic migration.

A communication graph must have at least one free reception/transfer window
per cyclic pipeline or a proven initial token.
The compiler checks capacities, blocking WAITs, and
join orders per channel. General deadlock freedom of arbitrary programs
is undecidable; unknown cases receive no proof status.

Static assignment is the real-time standard. Dynamic selection of a
free application core is allowed for tasks without hard deadlines, but their
waiting time must then not be considered statically proven.
The four system cores remain reserved for dispatcher/I/O.

### 28.6 Local ABI

| Register | Role |
|---|---|
| R0..R3 | parameters, R0 return value; caller-saved |
| R4..R5 | temporary; caller-saved |
| R6..R9 | callee-saved |
| R10 | local stack pointer |
| R11 | link address; caller must save it before a nested call |
| ACC | caller-saved |
| S0..S3 | task bindings, unchanged by normal fn calls |

The stack grows downward from a DATA-slot upper bound defined by the linker
and is 4-byte-aligned. The linker determines the maximum
from the call graph, spill requirements, and declared recursion depth.
ST/LD check stack accesses just like all local data.

Four parameters are passed in R0..R3, further ones in a
static argument block. Values above 32 bits use explicit
low/high pairs or a local result block; no hidden
128-bit register state is assumed.
External data is resources plus offset/length, not local pointers.

A code image may occupy several CODE slots. Functions within
the same image use JMP link forms. An overlay change is a
dispatcher/task change, not a normal call through a DMA-owned slot.
The ABI is identical for assembler, compiler, and future app.

### 28.7 Libraries and Numeric Processing

Software division uses one shared bounded 32-step
algorithm for unsigned quotient/remainder. Signed division handles
signs separately; rounding is toward zero. Division by zero is
a task error. INT_MIN/-1 returns INT_MIN under the word modulo contract;
this behavior must be explicit in the backend report.

CTZ can be expressed through bit isolation and CLZ:
if x=0 ->32, otherwise 31-CLZ(x AND (0-x)).
ABS/MIN/MAX/CLAMP use CMP, SEL, and SUB; signed INT_MIN remains
INT_MIN under modulo ABS, while saturating ABS is an explicit library choice.

Fixed-point library formats carry width, sign, and fraction bits in their
declared library contract; their source values are word or bounded word arrays.
Multiplication uses MUL-low/high or ACC;
rounding, shift, and saturation are explicit operations.
Reciprocal/root tables and Newton steps are approximations with
documented input ranges and error budgets, not exact division.

Decoder bit readers, byte order, and bounds checks form shared
compiler/guest library components. Creating a uRISC version
follows the same format semantics and golden vectors of the active modules;
no new incompatible bitstream is introduced.

### 28.8 Slot Planning and Ping-Pong

Example for a small streaming-capable task:

| Slots | Use |
|---|---|
| 0..2 | CODE, at most 3 KiB for the phase |
| 3 | stack, descriptors, small state |
| 4 | input A |
| 5 | input B |
| 6 | output A |
| 7 | output B |
| 8..11 | table/history/working data per phase |

Descriptor and data bank are separate. During RREAD into input B,
EX works on input A; after completion, the roles switch.
RWRITE from output A permits computation into output B.
This allocation is an example, not a guarantee for every decoder.

The slot planner tracks live intervals and DMA leases per bank.
Spilling to an active DMA slot is prohibited. If the working set does
not fit, it reduces the phase, places state in a RAM resource,
or rejects the target profile with required/free size.
It cannot silently turn a local LD into a RAM access.

### 28.9 Budget Calculation and Admissibility

A task budget covers:
pipeline startup/drain, issued instructions, two windows per
taken branch, software libraries, message attempts,
transfer waiting time, overlays, and dispatcher work.

For a finite trace without WAIT, the control value
`cycles = issued + 3 + 2*takenBranches` must be used, provided
pipeline startup/drain and branch windows are not double-counted.
For multiple sections, boundaries are determined from the event model;
a blanket formula does not replace the trace.

RUN: 4 ns per core cycle, IDLE: 20 ns.
A 60-Hz frame lasts 16.666666... ms, mathematically about 4.166 million
RUN cycles per core. Four cores supply about 16.666 million
issue windows per frame before control/wait overhead.
60 application cores supply a theoretical maximum of 15 GInstr./s,
all 64 together 16 GInstr./s.

A PCM block of 1024 samples per channel at 48 kHz lasts
21.333333... ms: about 1.066 million IDLE cycles.
At 192 kHz, it is only 5.333333... ms or 266666 IDLE cycles.
An RAU IDLE validation must name rate, channels, and profile.

A hard-deadline report requires bounded loops, admissible
worst-case I/O pauses, bounded retry/backpressure, and buffered
media starts. If a bound is missing, the result is
UNPROVEN with its cause, not PASS.

### 28.10 Diagnostics and Build Product

Each rejection states source location, task, affected resource,
calculated requirements, and target limit. Examples:

```text
task decode_stereo: local working set 28672 bytes, available 12288
loop reconstruct: dependency on previous sample, not parallelizable
task video_root: deadline not proven; reference RAM pause unbounded
spawn workers: at most 72 simultaneously active tasks, target permits 60
```

A deterministic build contains:
ISA ID, target profile ID, source/compiler hash, input resource hashes,
CODE/DATA images, relocations, entry PCs, task graph, slot/ownership plan,
resource rights, cycle model version, debug source mapping, and report.
Identical sources, versions, and options produce identical guest bytes.
Timestamps and host paths belong in separate provenance metadata.


### 28.11 Ownership, Channels, and Task Errors

A read-only borrow is task-local and may not escape to another task.
Physically shared read-only RAM is a separate immutable region capability:
each reader has a validated binding and no writer may exist during its lease.
A mutable borrow exclusively owns its entire declared interval.
Send reservation immediately excludes the sender from use, mutation or a second
move. A successful visible transfer commits revocation/generation reassignment
and gives the receiver the new binding. Failure restores the sender's ownership
before language unwind. A DMA completion alone is not an endpoint move.
Local addresses are core-specific and cannot be moved as pointers
into another core context.

A channel has static element size and capacity.
32-bit scalars fit directly in USER. Larger elements reside in
a bounded RAM ring; USER transmits slot/sequence as a token.
The compiler runtime validates token, generation, and element size.
Queue capacity is the declared channel capacity, not
automatically the Rx FIFO size.
The source channel consists of distinct move-only send/receive endpoints;
USER carries a token, never an unvalidated shared source owner. Direction,
element type, closed endpoint state and channel closed/drained state are checked
separately. Capacity zero is rendezvous; the runtime must reserve bounded
pending-send/receive state and must not misrepresent hardware FIFO buffering
as a language message queue. No null messages are transported.

Smallsome Code send may wait according to the language contract; its lowering consists
of nonblocking SEND/status checking plus explicit WAIT.
Smallsome Code receive is correspondingly a RECV check plus a WAIT loop.
Rx may also contain completion/service messages: the runtime
demultiplexes by kind and source into bounded task-state regions.
It must neither discard another recipient's message nor arbitrarily often
reinsert it into the same hardware FIFO.

A task error invalidates its output regions only after
orderly cancellation. Successors receive ERROR instead of successful ownership.
race/stop terminates the loser cooperatively; already visible external
writes are not undone. The ownership plan must therefore
separate loser outputs from regions in active use.

Race selects the earliest successful completion tick; equal ticks use the
lowest one-based SOURCE INPUT index, not coreID or Rx arrival order. The event
model's coreID order does not override this language tie-breaker. A completion
at or before the deadline precedes timeout, even if the dispatcher observes it
later. Later completions cannot win over an expired deadline. Source timestamps,
input indices and delivery status must fit the bounded dispatcher descriptor.

Cancellation is signaled immediately but a language scope does not finish
until children, timers, Tx and DMA leases reach terminal states. External
effects already visible are never rolled back. Every admissible hard-deadline
program includes cancellation/drain WCET and device stall bounds; otherwise
the result is UNPROVEN. A waiting task does not automatically release its
core/local slots. Releasing execution capacity requires an explicit dispatcher
safe point and resource plan; no implicit scheduler migration is introduced.

Returning word/scalar, tuples or owned resources protects the result region
before RAII; caller R0 or a bounded result block receives it only after cleanup.
Guru cannot access released locals. Structured Error/Result records use a
bounded ABI representation including state, numeric code and source location.
Language error codes (e.g. 1102 ArithmeticError) and hardware fault IDs (1..10)
remain distinct fields. Truncated diagnostic text must be flagged, not silently
represented as a complete host Error. Hardware FAULT requires dispatcher action
and cannot resume inside a source Guru handler.

### 28.12 Assembly Syntax and Concrete Examples

Canonical modes are constructed from the tables in section 11.
The dotted names are fixed assembler words:

| Instruction | Mode words |
|---|---|
| ADD/SUB | scalar, lane16, lane8; wrap, usat, ssat |
| CMP | scalar, lane16, lane8; eq, ne, slt, sle, ult, ule |
| SHL | scalar, lane16, lane8; imm or reg |
| SHR | scalar, lane16, lane8; logical or arithmetic; imm or reg |
| MUL | unsigned or signed; low or high |
| MAC | unsigned or signed; add or set; readlow or readhigh |
| LD | u8, u16, u32, s8, s16 |
| ST | u8, u16, u32 |
| JMP | rel, indirect, linkrel, linkindirect |
| WAIT | message, ticksimm, ticksreg |

Mode names with multiple parts are joined by dots in table order,
for example `lane8.usat` or
`scalar.arithmetic.reg`. Inapplicable parts are rejected.
For LD/ST, the mode is attached to the mnemonic, for example `LD.u16`.
Register counts in shift instructions appear as Rc instead of an
immediate value. JMP.rel/JMP.linkrel use a label or I22,
JMP.indirect/JMP.linkindirect exactly one register.
WAIT.message has no operand; WAIT.ticksimm U22,
WAIT.ticksreg lowRegister,highRegister.
WAITUNTIL lowRegister,highRegister remains a canonical mnemonic.
MAC.readlow/readhigh has only Ra; the assembler sets B/C and
the signed bit to zero. All other MAC forms have three registers.

Lexical contract:
ASCII mnemonics, case-insensitive mnemonics,
case-sensitive labels, R0..R11/S0..S3, comma as operand separator,
`#` as comment, hex values with 0x, optional minus before numbers.
No implicit octal interpretation. One statement per line.

Label addresses are byte addresses; the assembler calculates
relative word distances from them. `.word` writes u32 little-endian,
`.byte` exactly the checked u8 values. `.align n` is allowed for
powers of two up to 1024 and fills with zero.
CODE padding uses NOP words instead of opcode 0.

Example A, scalar sum 1..4, slot 0 CODE and slot 3 DATA.
R3 is a deliberately constructed zero-valued register; R0 remains writable:

```text
.code 0
start:
    MOVI R0, 0
    MOVI R1, 1
    MOVI R2, 5
    MOVI R3, 0
again:
    ADD R0, R0, R1, scalar.wrap
    ADDI R1, R1, 1
    BNE R1, R2, again
    MOVI R4, 3072
    ST.u32 R0, [R4 + 0]
    DONE
```

Expectation: R0=10, DATA[3072..3075]=0A 00 00 00, three taken
BNE and one untaken BNE, STOPPED without open transfers.
19 issued instructions plus 3 pipeline windows and
6 branch windows yield 28 RUN cycles = 112 ns in the basic model.
DONE drain is included in the shared pipeline drain here.
A later trace must reproduce this specific reference.

Example B, branchless maximum of a signed value and zero:

```text
    MOVI R0, 0
    CMP R3, R1, R0, scalar.slt
    SEL R3, R0, R1
```

R3 is 0 for negative R1, otherwise the unchanged value of R1.
The result also handles INT_MIN correctly, without ABS overflow.

Example C, data-bank RREAD and cookie:
R4 points to a validated descriptor in slot 3,
whose data destination is slot 4 and whose cookie is 17.
S0 denotes the readable data window.
R9 contains zero; R6 is only the status destination.

```text
request:
    RREAD R6, S0, R4
    BNE R6, R9, request_failed
poll:
    RECV R6, R7, R8
    BEQ R6, R9, dispatch_event
    WAIT.message
    JMP.rel poll
```

dispatch_event checks kind/source in R8 and cookie in R7;
READ_DONE with status OK makes slot 4 usable.
Other events go to the designated bounded demultiplexer.
request_failed handles BUSY/errors; immediate infinite retry
is not an admissible real-time plan. The two target labels are
application continuations, not hidden hardware instructions.

### 28.13 Build Manifest and Theoretical Programs

The build product consists of a canonical UTF-8 JSON manifest
plus referenced CODE/DATA/resource bytes. It is a program
description and changes none of the Generation-26 bitstreams.
WORM0 can transport the required individual resources as before;
no second universal media container is defined.

Binding manifest fields:

| Field | Meaning |
|---|---|
| schema | urisc-t1-program-2; version 1 manifests lack the audited language/runtime contract |
| isa | urisc-t1-1.0 |
| model | specific event model version |
| profile | 64-core standard profile and explicit latency/error parameters |
| language | Smallsome 3.0 revision, modes by canonical source ID, num/word capabilities, pinned Unicode version |
| inputs | IDs, lengths, SHA-256, and format generation |
| tasks | IDs, core mask, entry, code/data images, ABI parameters |
| slots | twelve roles per task, initialization bytes, and leases |
| resources | logical ID, generation, size, rights, traffic class |
| edges | producer, consumer, region, ownership/event condition |
| budgets | deadline, instruction/transfer limits, and proof status |
| runtime | contract urisc-t1-runtime-2, task/queue/timer caps, bounded Error/Result layouts, cancellation/drain limits |
| debug | source mapping and named breakpoints |
| expected | output hashes, final states, and expected errors |

Integers are stored in JSON as decimal numbers; 64-bit time and
40-bit addresses as decimal strings so that a JSON reader with
IEEE-754 numbers does not round them. Binary images are little-endian.
Manifest lists have a fixed order by ID, fields fixed names;
hash objects are not identified by a host path.

The loader checks schema, ISA, latency profile, lengths and hashes,
disjoint rights, initialized code/data regions, entry alignment,
and admissible system/application cores before changing state.
An erroneous manifest is rejected as a whole.
Importing v4.1 requires reassembly for T1, not
automatic interpretation of old opcode bytes.

## 29. Smallsome Formats on uRISC (Normative Integration Contract)

### 29.1 Source, Generation, and Host/Guest Boundary

All eight format families use Generation 26. Their binary
grammar in the active modules and release contracts
under [Formats](../../Formats/README.md) remains authoritative.
This chapter specifies their execution on T1 and does not copy
a second divergent file-format standard.

A theoretical decoder is a T1 guest program made of normal opcodes.
The native Retro decoder supplies golden results and metadata for
comparison. Calling it is not a simulated core task and must
account for no uRISC cycles or core counts.
A decoder not yet ported is marked HOST_REFERENCE/UNPORTED.
The future app may use it for preview;
no guest performance report is then provided.

Codec engine logic knows only memory. Host files, preview, audio output,
and Scrapbook persistence use the existing project boundaries.
Audio goes through the central mixer, images through central
image composition. Containers delegate embedded codecs; they
do not implement RAU/RFXL a second time.

### 29.2 Format Matrix and Task Boundaries

| Family | Active contract | T1 work unit | Dependency |
|---|---|---|---|
| RAU / RC26 | mono/stereo PCM | sample section within a frame or RC packet | bit reader, predictor/LMS, possibly previous frame |
| RFXL | RGB24/RGBA32 image | existing stream, then line/tile | LZSS, MTF, and spatial predictor |
| RFXA / RFXZ | animation, optional RAU sprites | frame/delta; decompressed group | persistent canvas and palette |
| PMF0 / PMFZ | video, optionally continuous RAU | existing root per frame; group for PMFZ | history frames, copy-current within a root |
| RPMC | song, patterns, sample/instrument state | validated loading sections and time events | pattern/voice/mixer state |
| LeXA | UTF-8 and bitmap-font resources | existing text/font block | compression stream and glyph boundaries |
| FormA / FMA1 | Geometry-Core and Scene | geometry/scene loading phase, later render tile | delta values, indices, materials |
| WORM0 | single typed resource | existing resource chunk | chunk decoder, then embedded codec |

A 1-KiB slot is a transport/working bank, not a new
bitstream block format. 16..64-KiB blocks are not retrospectively
invented as a universal guarantee of independence.
Byte and bitstream boundaries remain exactly preserved.

### 29.3 RAU-26 and RC26

Sources: [RAU module](../../Formats/module_format_rau.pbi),
[release audit](../../Formats/RAU_RFXL_RELEASE_AUDIT.md).
RA26 has a 17-byte header with flags, explicit rate, and
sample count; frames contain up to 1024 samples per channel.
The current decoder reconstructs fixed predictor/Rice, optional LMS,
constant frames, and copies of the previous frame.
A current RA26 file is therefore not generally frame-parallel.

RC26 contains packets with validated, matching metadata.
Each embedded RA26 packet starts its own decoder state;
packet parallelism is possible after header/length validation.
Within a packet, decoder dependencies remain.
The final PCM join preserves sample order.

Native reference output is signed PCM16 interleaved, with sample counts per
channel. The format supports mono/stereo; 7.1 results from mixing multiple
sources or other PCM resources, not from eight-channel RA26.

The existing stereo decoder already uses:
history 2*1024*2*4=16384 bytes, residuals 8192 bytes, and
transform buffers 8192 bytes, totaling 32768 bytes before code/additional state.
Unchanged, it does not fit in 12 KiB locally.
The T1 plan uses RAM for history and bounded local
sample sections; an optimized streaming liveness plan must demonstrate
correct predictor/LMS order.
These are porting requirements, not a claim that a
12-KiB implementation already exists.

Data flow:
RAM input ring -> header/bit reader -> residual section ->
predictor/LMS -> inverse stereo transform -> PCM16 ->
shared audio input ring -> central mixer -> final output ring.
DMA completions protect slot changes.

Suitable opcodes: LD.u8/u16/u32, SHR/SHL, AND, CLZ for
bit-reader helpers, ADD/SUB, signed MUL/MAC, CMP/SEL, and ST.
Rice escape is bounded by the format; nevertheless, bit-reader boundaries are
validated before every refill. 64-bit intermediate values are represented
through register pairs or ACC.

The validation case must cover lossless and all lossy profiles,
mono/stereo, final short frames, previous-frame mode, RC packet boundaries,
seek, and erroneous/truncated streams.
The hypothesis of one core at 50 MHz is confirmed only after T1 instruction
and worst-case transfer measurements for a specified rate.

### 29.4 RFXL-26

Source: [RFXL module](../../Formats/module_format_rfxl.pbi).
Output is RGB24 or RGBA32; alpha is reconstructed according to the existing
format. Palette, Planar, Raster, and Sparse-Screen remain
existing modes. Crappy/Low/Mid/High/Lossless are encoder profiles,
not new uRISC decoder opcodes.

Codec paths include raw, RLE, BytePack, Rice, and LZSS.
LZSS has a 12-bit window; back-references may overlap
and must be reconstructed in a defined copying order.
A Memcpy for nonoverlapping regions does not replace this case.
Predictors require previous pixels/lines; MTF has a
continuously mutated palette-index order. Such streams are serial
until a proven reset/substream boundary is reached.

A 720p RGBA image has 3686400 bytes. One RGBA line has 5120 bytes;
two lines already occupy 10240 bytes. Together with 4096 bytes of
LZSS history, palette, code, and stack, this does not fit locally.
Decompressed streams, history, or lines therefore reside in RAM;
local sections and DMA windows are planned individually.
A planner must check T1 throughput including these transfers.

Data flow:
stream validation -> serial decompression/unpredict where required ->
pixel/palette reconstruction -> RGBA RAM ->
parallel composition on disjoint tiles -> framebuffer.
Parallelizing encoder candidate search does not
parallelize the same decoder bitstream.

PERM sorts bytes, CMP/SEL produce masks, lane ADD/SUB support
provably suitable pixel arithmetic. Signed intermediate values for MED and
other predictor rules must not be distorted by unsigned lane saturation.

A single image has no intrinsic 60-Hz deadline.
The application specifies image size, loading/presentation deadline, and
permitted decoding phases. A universal statement that RFXL runs at IDLE in
real time without an image size/deadline is invalid.

### 29.5 RFXA-26 and RFXZ-26

Source: [animation module](../../Formats/module_format_rfxlanim.pbi).
Full frames use RFXL; delta frames modify a persistent
canvas. Existing global palettes, rectangles, and 16x16 delta regions
remain the format boundaries. Canvas decoder order is binding.

After validation, a delta may be distributed across multiple cores
only if its write regions are disjoint or the original
ordering effect is preserved. A later frame starts only
after completion of the previous canvas state.
Optional RAU frame chunks are audio sprites; they must be distinguished from
PMF0's continuous audio track.

RFXZ must decompress the required existing group.
Group size and peak RAM are checked before starting.
Neither wrapper nor root index eliminates decompression costs.
The canvas is marked presentable only after a complete frame.

### 29.6 PMF0-26 and PMFZ-26

Sources: [video module](../../Formats/module_format_pmf0.pbi),
[release audit](../../Formats/PMF0_RELEASE_AUDIT.md).
PMF0 uses existing frame/GOP/root indices, quadtree leaves,
Copy-current, Copy-reference, Reference-Residual, Motion, Flat,
BiColor, and QuadColor. RGB/YCoCg and prepared chroma profiles are
inversely transformed exactly according to active Generation 26.

Frame history may use lags 1,2,4,8. A reference is immutable until
its final reader finishes. At 720p, eight RGB24 history images occupy
8*1280*720*3=22118400 bytes (21.094 MiB), plus the current RGB image
at 2.637 MiB and RGBA framebuffer at 3.516 MiB.
These values fit in 256 MiB, but not in local slots.
Further wrapper, bitstream, audio, and rendering buffers are additional.

Before a frame:
validate indices, root geometry, stream boundaries, and reference lags;
then distribute independent roots to free application cores.
The active encoder limits Copy-current to the current root.
The decoder, by contrast, checks preceding image positions in particular;
an externally supplied stream must therefore not be treated as
root-independent without validation. The guest validator checks Copy-current
source regions and creates dependency edges for cross-root accesses
or decodes the affected section serially.
Within a root, quadtree/bit-reader order and Copy-current
dependencies remain; a compiler must not split these pixel by pixel.
It must check the Copy-current source region against the already reconstructed
region under the existing decoder contract.

A root with 64x64 RGB requires 12288 bytes for pixels alone;
with code and state, it does not fit entirely locally.
The guest works in subsections and RAM history without turning these into
new independent bitstream roots.
Even a 32x32 root requires 3072 bytes of RGB plus state/reference windows.

Frame barrier:
all roots successful -> color/alpha/composition phase ->
all writes visible -> PRESENT.
Frame n+1 may use references from n only after this reconstruction barrier.
Audio is forwarded independently according to sample time.

PMF0 audio is continuous RAU-26/RC26 with fragments between
video frames, not a sequence of audio sprites. Fragment boundaries are not
automatically codec reset points. The demuxer feeds bytes into the
existing RAU streaming state.

PMFZ groups independent PMF0 units and optionally uses
BriefLZ for an entire group. Decompression requires the whole
group block beforehand; startup latency and peak RAM must be budgeted explicitly.
Documented encoder practice may choose plain PMF0 for audio.
No lab may account for PMFZ memory or decompression time as zero.

Four RUN cores for 720p60 theoretically correspond to about
16.666 million issue windows per frame, about 18.08 per pixel before
branches, bit reader, reference transfers, and synchronization.
This is a hypothesis to be checked, not decoding evidence.
Known ARM/PureBasic timings are not converted to uRISC cycles.

### 29.7 RPMC-26

Source: [music module](../../Formats/module_format_rpmc.pbi).
RPMC is loaded as a validated song with orders, patterns, events, samples,
and instrument/surround state. It is not a music bitstream
to be decoded anew for every audio frame.

The runtime separates time events from the central sample mixer:
pattern/tick service -> voice parameters -> existing sample data ->
mixer sections -> final PCM ring. Sample/instrument slots remain
semantically sparse according to RPMC, not artificially densely duplicated.
Effect state, order changes, and tempo changes advance serially
according to music semantics.

Voice computation may be distributed into disjoint local partial results.
The join mixes in a fixed order with defined
width/rounding. 7.1 assignment is mixer/channel state, not a new
RPMC or RAU opcode.

### 29.8 LeXA-26

Source: [text/font module](../../Formats/module_format_lexa.pbi).
Text resources supply UTF-8 bytes and existing indices;
bitmap-font resources supply validated geometry and glyph blocks.
Unicode decoding, glyph selection, and screen rasterization are
application/font tasks. TTF/OTF import belongs to preparation,
not to an invented LeXA runtime parser.

Guest sequence:
header/lengths -> existing compression path -> text/glyph RAM ->
layout -> disjoint text tiles -> central image composition.
Compression streams are parallelized only at actual resource boundaries.
1bpp glyphs use LD.u8, masks, and SEL; row/column bit arrangement
remains that of the existing font contract.

### 29.9 FormA-26 / FMA1

Source: [geometry module](../../Formats/module_format_forma.pbi).
Int32-mm coordinates, triangle indices, delta-varint/byteplanes/PackBits,
scene attributes, and embedded RFXL textures remain unchanged.
Geometry-Core and the Scene appendix do not duplicate geometry.

Data flow:
scene validation -> geometry/attribute decoding -> geometry RAM ->
transform/clip -> tile lists -> tile renderer -> framebuffer.
Delta/varint order remains binding in the loading path.
Rendering is a subsequent use, not a property of the codec.

Fixed-point transformation, perspective, and depth specify value ranges,
rounding, and error budget. A 16-bit Z buffer is a scene decision,
not a general guarantee for Int32-mm geometry.
Texture RFXL is loaded through the same codec path. Texture tiles
are made available locally through DMA before use.

### 29.10 WORM0-26

Source: [resource container](../../Formats/module_format_worm0.pbi).
WORM0 contains exactly one typed resource and the existing
chunk boundaries/methods. The guest validates the overall payload, decompresses
existing chunks, and delegates to RAU,
RFXL, LeXA, FormA, or RAW according to resource type.

Independent existing chunks may be distributed across multiple cores
where their decoder state starts anew for each chunk.
A decoded payload must reach exactly the declared size.
A decompressed but semantically invalid embedded resource
remains an error, not a decoder-side RAW fallback.

Raw-data fallback is an existing encoder decision; it must not
conceal corrupt codec data.
Compression and embedded codec are timed separately,
then added as a complete loading path.

### 29.11 Porting, Equality, and Measurement Protocol

The same validation sequence applies to every family:

1. Record the current native golden corpus and module hash.
2. Capture binary boundaries and numeric intermediate values from the active
   decoder; do not port a historical format version.
3. Develop the same algorithms as T1 guest code or compiler
   lowering and assemble them for the specified slot plan.
4. Check output against the native reference: lossless byte-exact,
   lossy also byte-exact against the decoder of the same bitstream.
5. Check EOF, truncation, erroneous lengths/offsets, boundary dimensions,
   and resource failure with defined errors.
6. Log instructions, cycles, transfers, peak RAM, peak slots,
   messages, underflows, and deadline violations.

A codec profile is UNPORTED, FUNCTIONAL, TIMED, or BOUNDED.
FUNCTIONAL requires output equality; TIMED a modeled trace;
BOUNDED additionally requires the documented worst-case assumptions.
One passing clip does not make the entire format family BOUNDED.

## 30. Theoretical Lab and Scrapbook Contract (Specification)

### 30.1 Scope of the Future App

The app will be built according to this specification. It executes T1 code,
visualizes an 8x8 CPUlet, and saves reproducible experiments.
RTL, FPGA, physical manufacturing, and automatic tool installation
are not prerequisites for this theoretical project.

The current PureBasic emulator is consolidated into a single T1 reference core.
GUI, assembler, compiler, execution display, and headless
verification use the same core. Format modules remain native
references, and the existing memory contracts remain in place.

### 30.2 Event Time and Determinism

A T1 event model uses integer system ticks.
RUN issues every two fabric ticks, IDLE every ten.
The model order at a point in time is:
completed external writes/DMA -> message/timer wakes ->
core issue windows in coreID order ->
fabric arbitration -> scheduling of new device events.
Newly accepted requests respond no earlier than a later tick.
A FIFO entry accepted only at the end of a tick cannot
be read retroactively in the same core window.

All round-robin pointers start at source 0. For simultaneous
events with equal priority, the stored
(tick, event phase, source ID, local sequence) order decides.
Quotas are nonnegative tick credits per class. At
100-tick boundaries, 4/12/64/20 credits are added; at most one
full quota is additionally carried into the next section.
Each transmitted beat consumes one credit. Borrowing free credits
is allowed only when the original class has no waiting beat.
For RAM service, five credits are reserved at the start; this also makes
the service window, which cannot be interrupted beat by beat, deterministic.

Long idle phases are skipped up to the next event
without changing guest cycles or deadlines.
Host parallelization must not change the specified guest order.
Every random/error generator has a saved seed.
Host audio/image preview follows guest state and does not control its time.

### 30.3 Experiment Contents

A Scrapbook entry contains:
unique experiment ID, title/note, source code, ISA/compiler/model version,
64-core target profile, resources and their SHA-256 hashes,
initial state, seed, event/latency profile, expectations,
execution result, measurements, and proof status.
Comparable repetitions refer to the same input baseline;
changed sources/profiles produce a new experiment version.

A canonical snapshot contains PC, registers, ACC, slots/roles,
resources, FIFO/Tx/tags, timers, RAM, dispatcher, and all scheduled
external events. Saving only PC/registers is insufficient
for continuation or backward stepping.

The app may save snapshot deltas; it must reconstruct the same
complete state from them. A checksum verifies it.
Cancellation, fault, timeout, and PASS are separate results.
Actual Mac runtime and simulated guest duration are displayed separately.

### 30.4 Observation and Control Contract

The 8x8 grid shows ID, state, clock level, task, and activity.
Selection shows PC/instruction, registers, ACC, slot roles/ownership,
messages, and open transfers. A timeline shows core
sections, fabric occupancy, DMA, deadline, and present/audio events.

Instruction single-step means: run until the next completed
instruction of the selected core; the other masters continue
according to their guest time. System-tick single-step
is a separate operation. A halt freezes the entire model state.

Breakpoints trigger before IF of the target instruction; watchpoints report after
a valid completed write. A debug halt requires no
guest FAULT and changes no queue.
Backward stepping is performed through snapshot plus deterministic
replay, not through invented inverse opcodes.

Experiments for RAM pauses, full FIFOs, USB stalls, and underflow
are saved as profiles. The app then determines which
budget assumption was violated instead of interpreting host stutter as
a uRISC performance problem.

### 30.5 Limits and Future Delivery

The app may display graphics/audio output without physical PHYs.
The display proves neither TMDS/USB/PDM compliance nor
actual ASIC clock capability.
Code signing protects the origin/integrity of the Mac app;
it certifies no architectural claim.
Developer ID, notarization, and bundle integration are implemented during
app construction using the existing Mac packaging paths.


### 30.6 Reproducible Model Profiles

The standard profile is named `t1-lab-nominal-1`. It is a synthetic
working assumption for experiments, not a reproduction of a measured LPDDR4 PHY:

| Model component | Nominal contract |
|---|---|
| Core clocks | RUN every 2, IDLE every 10 system ticks |
| Start/frequency change | only at tick boundaries; first issue at the next suitable global 2-/10-tick grid point |
| Group/router link | 1 beat per tick; quotas from 27.8 |
| Local message path | same acceptance/ACK contract; earliest acceptance in the tick after SEND |
| RAM base latency | packet ready for service 64 system ticks after complete acceptance |
| RAM data service | one shared unit; 5 ticks per packet up to 64 bytes, reads and writes together |
| RAM order | arbitration by traffic class/quota, FIFO within the same class |
| Refresh/bank conflicts | disabled in the nominal profile and identified as such in the report |
| Service command | earliest processing 1 tick after complete acceptance |
| Display | present at the next rational 60-Hz VBlank boundary |
| Audio | sample consumption at the configured rational rate |
| USB storage | synthetic 35000000 bytes/s, startup latency 500000 ticks, no random stall pause |
| Seed | 0 unless a random profile is active |

RAM read bytes are read at the end of data service; write bytes
are visibly written there. A service-ready packet waits for
the shared unit. Read responses then wait for the response link;
write ACK is generated only after the last visible packet.
The 5-tick service permits at most 6.4 GB/s of aggregate array payload
for full packets, never separate 6.4 GB/s for reads and writes.
A short transfer also occupies a service section.
A transaction data buffer is limited to eight requests per core plus
bounded autonomous device buffers; unbounded guest queues
are not assumed.

Frames and sample periods are rational numbers. The model uses
a phase accumulator; an event at a noninteger tick falls
on the first tick after the ideal time. This causes
at most one tick of quantization per event, no cumulative drift.
For example, VBlank intervals at 500 MHz/60 Hz follow a fixed
pattern of 8333333/8333334 ticks.

The stress profile `t1-lab-stress-1` uses the same rules and adds
every 3900 ticks a synthetic 100-tick RAM service pause.
A running service operation is paused and then resumed.
USB adds, after every 1048576 delivered bytes, 50000000 ticks of pause.
The numbers are deliberately saved test assumptions, not
guaranteed electrical DRAM/USB maxima.
Further profiles declare every deviation together with the seed.

A deadline met only in the nominal profile receives TIMED,
never automatically BOUNDED. A worst-case profile must demonstrate
upper bounds for its entire permitted input set.
Endless disruptions are logged as TIMEOUT through a bounded experiment
duration; they are not reinterpreted as successful runtime.

### 30.7 Research Experiments as the App's Baseline Set

| Experiment | Expected result / Question |
|---|---|
| ISA scalar | all boundary values and encodings; correct register/SRAM results |
| ISA lanes | no lane carries, correct masks and saturation |
| LD-use / branch | no stall; two flush windows only on a taken branch |
| 64-Core message ring | ordered tokens, no duplication, controlled FIFO load |
| Slot conflict | deliberate IF/EX/DMA conflict causes precise FAULT |
| Ping-Pong-DMA | read A while filling B; correct cookie and ownership |
| Shared RAM data | write completion before ownership transfer, then identical read |
| RAU stream | byte-identical PCM, rate/slot/deadline report |
| RFXL image | byte-identical RGBA for all existing decoder paths |
| RFXA canvas | ordered deltas and optional audio sprites |
| PMF0 roots | identical 1/4/60-worker output, correct history/root dependency |
| RPMC music | same event sequence, deterministic mixer join |
| LeXA / FormA / WORM0 | byte-identical resources, correct boundaries and nested codecs |
| Media combination | PMF0 plus audio plus input/USB; no claims that costs have disappeared |
| Stress / Underflow | defined error and violated budget assumption visible |
| Snapshot replay | identical final state and identical guest trace |

This is the binding future experiment catalog, not an app feature
created or executed in this round.

## 31. Conformance, Open Physical Questions, and References

### 31.1 Binding Theoretical Release Scope

T1 version 1.0 is a theoretical specification with:
31 opcodes, fixed operand/mode bits, register/memory semantics,
branch/WAIT contract, messages, resources, ownership and visibility,
compiler target/ABI, all eight format profiles, and a reproducible
lab contract. Its completeness means defined target rules,
not an already working compiler or proven chip.

Mandatory future model checks:

| Area | Properties to check |
|---|---|
| Encoding | all 31 opcodes, reserved fields, register limits, immediate values |
| Arithmetic | signs, boundary values, lane isolation, saturation, MUL-high, ACC forwarding |
| Memory | alignment, slot boundaries, initialization, IF/EX/DMA conflict, no partial ST |
| Pipeline | RAW/WAW, LD-use, branch flush, precise FAULT, WAIT race |
| Communication | FIFO quotas, retry/dedup, ordering, completion reservation, delivery fault |
| Resources | rights/generation, overflow checking, cookies, visibility, cancellation |
| Compiler | types, alias/ownership plan, ABI, branch relocation, slot spills, UNPROVEN |
| Formats | active golden corpora, pixel/PCM equality, dependency, truncation |
| System | boot, 64-core time ordering, buffer switching, audio/display underflow |
| Replay | snapshot completeness, same seed, same hashes and guest traces |

This list is the future test contract. No smoke tests or codec programs
are executed in this documentation round.

### 31.2 What Is Theoretically Decided and Physically Open

Decided for T1:
no LOOP/CTZ/DIV, scalar MUL/MAC, atomic local descriptor acceptance,
0-byte CPUlet-SRAM in the standard profile, 8 Tx/tags/completion slots,
16-byte contact header, 64-core numbering, fixed ABI and model time.

Open for future manufacturing:
process/PDK, SRAM/register-file/128-bit descriptor bank macros,
250-MHz timing, 500-MHz fabric, DRAM timing/refresh, PLL/clock domains,
PHYs and signal integrity, power supply, package, area,
power and temperature. Single-cycle MUL and
LD/descriptor forwarding in particular require physical validation.
An FPGA with a different clock frequency can confirm functional rules,
but cannot prove the ASIC frequency.

The PPA figures in section 21 remain historical
guidance with unknown accuracy. A theoretical
simulation run will never turn them into watt/mm2 measurements.

### 31.3 Source Baseline and Responsibility

Baseline of this normative specification: 10 October 2026.
Local active sources take precedence over old ZIP codec samples.
The archives serve only as provenance of emulator behavior.

- [Smallsome Code language](../../Ozon/CYMPLE3.md) and
  [grammar](../../Ozon/cymple3.ebnf).
- [Format overview](../../Formats/README.md) and
  [measurement matrix](../../Formats/FORMAT_MATRIX.md).
- [RAU/RFXL audit](../../Formats/RAU_RFXL_RELEASE_AUDIT.md) and
  [PMF0 audit](../../Formats/PMF0_RELEASE_AUDIT.md).
- Golden corpora: `Formats/ReleaseCorpus/`,
  `Formats/PMF0ReleaseCorpus/`, `Formats/RemainingFormatReleaseCorpus/`.
- [Project architecture](../../ARCHITECTURE.md) for memory/engine-area boundaries.

SHA-256 of the source baseline read during preparation:

| Source | SHA-256 |
|---|---|
| RAU module | 56a16198483c4a8a10c5e59dc1e086cebc2429a062f201b6b9bc73cb7e3b1331 |
| RFXL module | 7ab9534ab3101b992cfcc056f603fa2244a04dd45c8708b3c64d24047f7355b9 |
| PMF0 module | e6577ce03ecdc9b805c178d1d3fb574315fcf1fbf429c349be6d0203dcc34a4c |
| RFXA module | 3946992cc0198acbc9504ba56d1b3d36727447dee4bdcf79d3047cd56f32396f |
| RPMC module | 65bb2fa7ab6e3226fa2333e59111858e0b7b576c560a96e0f19db3003a4c90d7 |
| LeXA module | d85b26c21b1448c3e4db1727656bbf27d0076b5a7991c7a4e3f73823ead68057 |
| FormA module | 32a0d687ae9f548d413e00237f380f246e712c6193211cc014b9aa6f5db4d72b |
| WORM0 module | 9383fe7d40114c11b571d1402064bf7f860064168e5d4469553fdfb286191939 |
| Smallsome Code frontend (historical name: Cymple) | 233482e27a73233531234b8775934e5ba4211eb6023205ed83c35452a2b4cbc3 |
| Smallsome Code Ozon backend (historical name: Cymple) | 9d0adabf69a6c8729a636506ef67c3870e82ec13cee2f004de7335eb825e8dc0 |
| v4.1 emulator archive | 092b60b187990d5d49fc9444c48937ad40deb095e686fe80a91a51bc2dbdd1f7 |

Hashes document a reading baseline, not passed verification.
When a module changes, its T1 compatibility is checked using the same
integration contract and updated golden vectors.

### 31.4 Version Policy

Changes to instruction semantics or encoding increase the ISA version and
require reassembly, a new compiler report, and new experiment
snapshots. Language changes version the language; runtime/ownership changes
version the runtime contract; mandatory manifest changes version the schema.
These IDs are independent and all appear in an experiment. Revision 4 retains
ISA 1.0 but uses program schema 2 and runtime contract 2. Version-1 manifests
require an explicit migration/report, never silent acceptance as equivalent.
Editorial clarification alone increases only the document revision.

A future app may display several old Scrapbook experiments,
but execute them only with clear version association. T1 has exactly one
reference core; historical archives are not incorporated as a second active
architectural path.
