| git.druid.rocks | index | druid520 | kaboom | docs/ | superpowers/ | plans/ | 2026-09-30-vmm-per-process-addrspace.md |
docs/superpowers/plans/2026-09-30-vmm-per-process-addrspace.md
# Per-Process Address Spaces Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Replace kaboom's two-fixed-window shared-memory model with real per-process page tables, so an arbitrary number of processes (bounded by physical memory, not a window count) can be resident at once.
**Architecture:** A new `vmm.nsc`/`vmm.nsh` owns a physical-frame bitmap allocator (seeded from a real e820/PVH memory map) and per-process page table construction (4KiB pages, built by re-splitting the existing 1GiB identity map: low-memory + `heap_colosseum` stay shared across every process, the `0x400000`-`0x1ffffff` range becomes private per-process). `exec.nsc`'s `sys_exec` switches CR3 around each `elf_call_entry` call instead of remapping a shared window. `paging.nsc` (the old scratch-pool/`exec_shwin` trick) is retired entirely.
**Tech Stack:** nsc/nscc (kaboom's own compiler), GNU `as`/`ld` for the small amount of real x86-64 asm this needs, qemu (via the `qemu_*` MCP tools) for all real testing — there is no unit-test framework in this codebase; every task's verification is a real compile + boot + serial interaction.
**Spec:** `docs/superpowers/specs/2026-09-30-vmm-per-process-addrspace-design.md`
## Global Constraints
- Ring0-only — no TSS, no user/kernel segments, no privilege transitions. Every new page table entry uses supervisor-only (U/S=0), matching `boot.s`'s own existing huge-page entries.
- Usable RAM is capped at the first 1GiB regardless of how much e820 reports (what `boot.s`'s identity map already covers). Growing the identity map past 1GiB is out of scope.
- No per-thread stacks — execution stays synchronous/blocking exactly as today; the single shared kernel stack (`boot.s`'s `stack_bottom`/`stack_top`) is never touched by this project.
- Every new/changed file follows `alloc.nsh`'s own established discipline: a `.nsh` declares functions and (only where genuinely needed) globals; a caller declares only the specific global it actually touches, inline — never a blanket include of everything a subsystem owns (this exact mistake caused a real boot hang earlier this session, per `alloc.nsh`'s own comment).
- Every allocator-style failure returns `(u64)0`/`(ptr)0` (`arena_alloc`'s established sentinel) — never partial success, never a silent wraparound.
- Build via (from repo root, after `export PATH="/root/work/nscc/out:$PATH"`): `sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh`, in that exact order every time (`mk/b.sh` wipes `out/`). Boot/test via the `qemu_boot`/`qemu_wait_serial`/`qemu_serial_send`/`qemu_serial`/`qemu_stop` MCP tools — `qemu_wait_serial(text="$")` confirms a real boot. Run `pkill -f qemu-system-x86_64` as its own isolated shell call if ever needed (chaining it with other commands truncates the rest of the command in this sandbox).
## Review Focus
- **A half-built address space on frame exhaustion mid-`PT_LOAD`-segment-mapping** — the spec requires every already-allocated frame for that attempt to be freed before `sys_exec` returns failure; a test must actually exhaust the pool mid-construction and confirm the free count afterward matches the count before the attempt, not just that the exec failed.
- **A destroyed process's frames actually returning to the allocator** — `vmm_destroy_addrspace` walking only the process's *private* pd entries (2-15) and never touching the shared ones (0-1, 16-23); a test must create+destroy a process and confirm `/int/vmm`'s free count returns to exactly its pre-create value, not just "some number came back."
- **Two sequential processes never observing each other's leftover memory** — a fresh process's private frames come from the bitmap in whatever order they were freed, so a naive implementation could hand a process a frame still containing the *previous* process's data; confirm this is either zeroed on allocation or that the PT_LOAD/.bss-zero path already covers every byte the process can address (the spec doesn't call out explicit frame zeroing, so this needs an explicit decision, not silence).
- **Deep `sh` nesting and shebang chains still hitting their documented limits correctly** (9 processes / 8 shebang levels) now that they go through real address spaces instead of the retired `exec_shwin` counter — the old counter-based limit is gone along with `exec_shwin_live`; confirm the *new* limit (frame exhaustion) kicks in at a consistent, still-reasonable depth, not silently much earlier or much later.
- **The boot-time `sh` (`kmain`'s own `sys_exec("sh", ...)` call, depth 0) restoring the correct CR3 on return** — this is the one call site with no `kaboom_proc_pml4` parent slot to restore from; a test must confirm `sh` returning (`exit`, respawned by `kmain`'s own loop) still leaves the kernel in a sane state for the *next* boot-time `sys_exec("sh", ...)` call, not just that the first one works.
---
### Task 1: Capture and verify the PVH `hvm_start_info` memory map (empirical probe)
**Files:**
- Modify: `src/boot/boot.s:86-89` (`_start`'s first two instructions)
- Modify: `src/kernel/kmain.nsc` (add a temporary boot-time dump, removed again at the end of this task once verified)
**Interfaces:**
- Produces: a saved 32-bit physical pointer (`start_info_ptr32`, a new `.bss` global in `boot.s`) that Task 2 reads to locate the real `hvm_start_info` structure. No other task depends on this one's temporary dump code.
The qemu `-kernel` PVH entry point receives a pointer to `struct hvm_start_info` in `%ebx` at `_start`, per the Xen PVH boot protocol — but `_start` currently clobbers `%esp` as its very first instruction without ever reading `%ebx`, so that pointer is lost today. This task preserves it and empirically confirms the real struct layout before Task 2 writes real parsing code against it (this codebase's own established standard: "confirmed working empirically before wiring it in for real" is the exact phrase `boot.s`'s own PVH comment already uses for the entry-note decision).
- [ ] **Step 1: Save `%ebx` before anything else touches registers**
In `src/boot/boot.s`, add a new `.bss` label right after the existing `pd:` block (before the stack-guard-gap comment, so it doesn't shift `stack_bottom`'s own alignment):
```asm
.align 4
start_info_ptr32:
.skip 4
```
Change `_start`'s first two real instructions (currently `cli` then `movl $stack_top, %esp`) to:
```asm
_start:
cli
movl %ebx, start_info_ptr32
movl $stack_top, %esp
```
- [ ] **Step 2: Add a temporary magic-value check to `kmain.nsc`**
Since `start_info_ptr32` is an asm `.bss` label, not an nsc-declared global, nsc can't reference it directly (same constraint `paging_asm.s`'s own comment already documents: "no way to reference a raw assembler label from another file directly, only through a real extern'd function"). Add a tiny accessor in a new file `src/boot/boot_asm.s`:
```asm
.section .text
.global get_start_info_ptr32
.type get_start_info_ptr32, @function
get_start_info_ptr32:
movl start_info_ptr32(%rip), %eax
ret
```
Add a self-contained decimal-print helper + the magic check to the very top of `kmain`, before `klog_init()`'s own first call (temporary — removed in Step 5 of this task):
```c
i8 dbgbuf[24];
ptr sip;
u64 magic;
u64 pos;
u64 digit;
u64 tmp;
sip = (ptr)(u64)get_start_info_ptr32();
magic = (u64)*sip & (u64)0xffffffff;
/* build magic's decimal digits back-to-front into dbgbuf, same
* digit-extraction algorithm virtfs.nsc's own virtfs_put_dec already
* uses, just writing forward into a local array instead of a caller's
* buffer -- this whole block is deleted at the end of this task, so
* it's fine for it to be self-contained rather than reuse put_dec
* directly (kmain.nsc doesn't link against virtfs.nsc). */
pos = 0;
tmp = magic;
if(tmp == (u64)0)
{
dbgbuf[0] = (i8)48;
pos = 1;
}
else
{
while(tmp > (u64)0)
{
digit = tmp % (u64)10;
dbgbuf[pos] = (i8)(48 + digit);
tmp = tmp / (u64)10;
pos = pos + (u64)1;
}
/* reverse in place */
tmp = 0;
while(tmp < pos / (u64)2)
{
digit = (u64)dbgbuf[tmp];
dbgbuf[tmp] = dbgbuf[pos - (u64)1 - tmp];
dbgbuf[pos - (u64)1 - tmp] = (i8)digit;
tmp = tmp + (u64)1;
}
}
dbgbuf[pos] = 0;
klog_write("kaboom: pvh magic (decimal, expect 862897528): ");
klog_write(&dbgbuf[0]);
klog_write("\n");
```
Add `"$as" $asflgs "$src/boot/boot_asm.s" -o "$out/boot_asm.o"` to `mk/b.sh` right after the existing `boot.o` line, and add `"$out/boot_asm.o"` to the final `"$ld"` invocation's object list (right after `"$out/boot.o"`).
- [ ] **Step 3: Build and boot**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot via `qemu_boot` (kernel=`out/kaboom.elf`, disk=`out/disk.img`, memory_mb=128), `qemu_wait_serial(text="$")`, then `qemu_serial_send("klogs\n")` and read the output.
- [ ] **Step 4: Verify the magic value against the expected `hvm_start_info` layout**
`klogs` should show `kaboom: pvh magic (decimal, expect 862897528): 862897528` — `862897528` is `0x336ec578`, the real Xen PVH `hvm_start_info` magic value, in decimal (this language has no hex literal/format support convenient enough for a one-off debug check, hence decimal). If it matches: the struct is laid out as documented below (Task 2 proceeds using the exact field offsets given there). If it does **not** match, this is a stop-and-reassess point — re-derive the real layout empirically (a fuller byte-by-byte dump is worth building at that point) before writing Task 2's parsing code; do not guess further.
Documented layout to check the dump against (offsets in bytes from `sip`):
```
0: u32 magic (expect 0x336ec578)
4: u32 version
8: u32 flags
12: u32 nr_modules
16: u64 modlist_paddr
24: u64 cmdline_paddr
32: u64 rsdp_paddr
40: u64 memmap_paddr (only present if version >= 1)
48: u32 memmap_entries (only present if version >= 1)
52: u32 reserved
```
and each entry in the array at `memmap_paddr` (`memmap_entries` of them):
```
0: u64 addr
8: u64 size
16: u32 type (1 = usable RAM, other values = reserved/unusable)
20: u32 reserved
```
- [ ] **Step 5: Remove the temporary dump**
Delete the debug block added to `kmain.nsc` in Step 2 (the accessor `get_start_info_ptr32`/`boot_asm.s` stays — Task 2 reuses it). Rebuild and confirm the boot log (`klogs`) is back to exactly its pre-Task-1 content.
- [ ] **Step 6: Commit**
```bash
git add src/boot/boot.s src/boot/boot_asm.s mk/b.sh
git commit -F- <<'EOF'
ADD: capture PVH start_info pointer, verify real memory-map layout
MOD: preserves %ebx (the PVH hvm_start_info pointer) before _start
touches any other register, via a new start_info_ptr32 .bss slot and
a tiny asm accessor. empirically confirmed against a real qemu boot
that the dumped structure matches Xen's documented hvm_start_info
layout (magic 0x336ec578) before any parsing code depends on it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 2: Parse the PVH memory map into a common e820-style array
**Files:**
- Create: `src/kernel/vmm.nsc`, `src/kernel/vmm.nsh`
- Modify: `mk/b.sh` (add `vmm.nsc` to the build)
**Interfaces:**
- Consumes: `get_start_info_ptr32()` (Task 1, `boot_asm.s`).
- Produces: `global i8 vmm_e820_buf[6144]` (256 entries × 24 bytes — a fixed, generous ceiling: real BIOS/PVH memory maps this session has ever seen are under 20 entries, 256 is headroom in the same spirit as the 10-slot proc table, not an arbitrary guess), `global u64 vmm_e820_count;`, and `void vmm_e820_add(u64 addr, u64 size, u32 type);` (appends one entry, silently drops any past 256 — same documented-truncation convention `exec.nsc`'s own argv cap uses) for Task 13 (the BIOS/dynamite path) to reuse verbatim. `void vmm_e820_from_pvh(void);` fills the array from the real PVH structure.
- [ ] **Step 1: Write `vmm.nsh`**
```c
/* vmm's public api -- see vmm.nsc. */
void vmm_e820_add(u64 addr, u64 size, u32 type);
void vmm_e820_from_pvh(void);
global i8 vmm_e820_buf[6144];
global u64 vmm_e820_count;
```
- [ ] **Step 2: Write `vmm.nsc`'s e820 ingestion**
```c
/*
* vmm: physical frame allocation and per-process page tables. this
* file's own first job is turning whatever the boot path handed us
* (pvh's hvm_start_info, or dynamite's own real int 15h e820 call --
* see vmm_e820_from_pvh/Task 13) into one common, flat array of
* (addr, size, type) entries -- 24 bytes each, matching the real xen
* hvm_memmap_table_entry layout exactly (empirically verified against
* a real qemu boot, see the note on Task 1), so the pvh path is a
* straight per-entry copy and the bios path just fills the same
* fields directly.
*/
u32 get_start_info_ptr32(void);
u8
vmm_byte_at(ptr base, u64 index)
{
ptr p;
u64 word;
p = base + index;
word = (u64)*p;
return (u8)(word & (u64)0xff);
}
u64
vmm_read_u64(ptr buf, u64 off)
{
u64 v;
u64 i;
v = 0;
i = 0;
while(i < (u64)8)
{
v = v | ((u64)vmm_byte_at(buf, off + i) << (i * (u64)8));
i = i + (u64)1;
}
return v;
}
u32
vmm_read_u32(ptr buf, u64 off)
{
u64 v;
u64 i;
v = 0;
i = 0;
while(i < (u64)4)
{
v = v | ((u64)vmm_byte_at(buf, off + i) << (i * (u64)8));
i = i + (u64)1;
}
return (u32)v;
}
global i8 vmm_e820_buf[6144];
global u64 vmm_e820_count;
global void
vmm_e820_add(u64 addr, u64 size, u32 type)
{
ptr slot;
if(vmm_e820_count >= (u64)256)
{
return; /* silently drop past the fixed ceiling, same
* documented-truncation convention exec.nsc's argv
* cap already uses */
}
slot = &vmm_e820_buf[0] + vmm_e820_count * (u64)24;
*slot = (i64)addr;
*(slot + (u64)8) = (i64)size;
*(slot + (u64)16) = (i64)(u64)type;
vmm_e820_count = vmm_e820_count + (u64)1;
}
global void
vmm_e820_from_pvh(void)
{
ptr sip;
u64 memmap_paddr;
u32 memmap_entries;
u32 version;
u64 i;
ptr entry;
u64 addr;
u64 size;
u32 type;
vmm_e820_count = 0;
sip = (ptr)(u64)get_start_info_ptr32();
version = vmm_read_u32(sip, (u64)4);
if(version < (u32)1)
{
return; /* no memmap field at all on a version-0 struct --
* leaves vmm_e820_count at 0, vmm_init (Task 3) halts
* cleanly on that same as any other detection failure */
}
memmap_paddr = vmm_read_u64(sip, (u64)40);
memmap_entries = (u32)vmm_read_u32(sip, (u64)48);
i = 0;
while(i < (u64)memmap_entries)
{
entry = (ptr)(memmap_paddr + i * (u64)24);
addr = vmm_read_u64(entry, (u64)0);
size = vmm_read_u64(entry, (u64)8);
type = (u32)vmm_read_u32(entry, (u64)16);
vmm_e820_add(addr, size, type);
i = i + (u64)1;
}
}
```
- [ ] **Step 3: Add `vmm.nsc` to the build**
In `mk/b.sh`, add `"$nscc" "$src/kernel/vmm.nsc" -o "$out/vmm.s"` right after the existing `alloc.nsc` line, and `"$as" $asflgs "$out/vmm.s" -o "$out/vmm.o"` right after the corresponding `alloc.o` line, and add `"$out/vmm.o"` to the final `"$ld"` object list (right after `"$out/alloc.o"`).
- [ ] **Step 4: Temporarily call it from `kmain` and verify via `klogs`**
Add a temporary call right after `klog_init()` in `kmain.nsc` (add `include "vmm.nsh";` to `kmain.nsc`'s existing include block first). Reuse Task 1 Step 2's decimal-print helper (`dbgbuf`/the digit-extraction+reverse loop) exactly as written there — copy that same block, parameterized by whatever value is being printed instead of `magic`:
```c
vmm_e820_from_pvh();
/* print vmm_e820_count using Task 1 Step 2's decimal-print code,
* verbatim, with magic replaced by vmm_e820_count and the klog_write
* prefix changed to "kaboom: e820 entries: " */
```
Then, for the first entry only (index 0, enough to sanity-check the parse without building a full loop in throwaway code): print `vmm_read_u64(&vmm_e820_buf[0], (u64)8)` (entry 0's `size` field) the same way, prefixed `"kaboom: e820[0] size: "`.
Build, boot, `klogs`, confirm the entry count is small and sane (single digits, not zero, not some huge garbage number), and entry 0's size is a plausible, non-zero byte count (order of magnitude matching qemu's `-m 128` — exact accounting isn't the point here, "a sane, non-zero, non-garbage number" is).
- [ ] **Step 5: Remove the temporary kmain dump**
Same as Task 1 Step 5 — delete the temporary block, keep `vmm.nsc`/`vmm.nsh` themselves. Rebuild, confirm `klogs` is back to its pre-Task-2 content.
- [ ] **Step 6: Commit**
```bash
git add src/kernel/vmm.nsc src/kernel/vmm.nsh mk/b.sh
git commit -F- <<'EOF'
ADD: vmm.nsc - parse pvh memory map into a common e820-style array
MOD: vmm_e820_from_pvh reads the real hvm_start_info memmap (Task 1
confirmed the real layout against a live boot) into a flat, fixed
256-entry array every later frame-allocator/bios-path task shares.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 3: Physical frame bitmap allocator
**Files:**
- Modify: `src/kernel/vmm.nsc`, `src/kernel/vmm.nsh`
- Modify: `src/kernel/kmain.nsc` (replace `paging_init()` call with `vmm_init()`, temporarily — both coexist until Task 11 deletes `paging.nsc`)
**Interfaces:**
- Consumes: `vmm_e820_buf`/`vmm_e820_count` (Task 2), `heap_colosseum_limit()` (`alloc.nsh`, already `0x3000000`).
- Produces: `u64 vmm_alloc_frame(void);` / `void vmm_free_frame(u64 addr);` (4KiB frames, `(u64)0` sentinel on exhaustion), `void vmm_init(void);`, `u64 vmm_frames_free(void);` (Task 4 needs this for `/int/vmm`).
- [ ] **Step 1: Add the bitmap and its accessors to `vmm.nsc`**
```c
include "alloc.nsh";
void halt(void);
/* one bit per 4KiB frame, up to 4GiB of real RAM (2^32/4096/8 =
* 131072 bytes) -- a real, generous ceiling in the same spirit as the
* 10-slot process table, not an arbitrary guess. usable RAM is
* further capped at the first 1GiB regardless of what e820 reports
* (see this project's own spec -- growing the identity map past 1GiB
* is a separate, later project), so in practice only the first 32768
* bytes of this bitmap (1GiB/4KiB/8) are ever touched. */
global i8 vmm_frame_bitmap[131072];
global u64 vmm_frame_hint; /* next index to start scanning from --
* pure optimization, never load-bearing:
* every alloc still verifies the bit is
* actually clear before taking it. */
u64
vmm_bit_get(u64 frame_idx)
{
u64 byte;
u64 bit;
byte = (u64)vmm_byte_at(&vmm_frame_bitmap[0], frame_idx / (u64)8);
bit = frame_idx & (u64)7;
return (byte >> bit) & (u64)1;
}
void
vmm_bit_set(u64 frame_idx, u64 val)
{
ptr slot;
u64 byteoff;
u64 bytebit;
u64 wordoff;
u64 byteinword;
u64 byteval;
u64 word;
/* *ptr is always a full 8-byte store here (no byte-granular write
* exists) -- same read-modify-write-the-whole-word technique
* proc_copy_name (exec.nsc:144-147) already uses: read the
* containing 8-byte word, clear/set only this one byte within it
* via shift+mask, write the whole word back. */
slot = &vmm_frame_bitmap[0];
byteoff = frame_idx / (u64)8;
bytebit = frame_idx & (u64)7;
wordoff = byteoff & ~(u64)7;
byteinword = byteoff & (u64)7;
word = (u64)*(slot + wordoff);
byteval = (word >> (byteinword * (u64)8)) & (u64)0xff;
if(val == (u64)1)
{
byteval = byteval | ((u64)1 << bytebit);
}
else
{
byteval = byteval & ~((u64)1 << bytebit);
}
word = word & ~((u64)0xff << (byteinword * (u64)8));
word = word | (byteval << (byteinword * (u64)8));
*(slot + wordoff) = (i64)word;
}
```
```c
global u64
vmm_alloc_frame(void)
{
u64 total_frames;
u64 i;
u64 idx;
total_frames = (u64)0x40000000 / (u64)4096; /* capped at 1gib */
i = 0;
while(i < total_frames)
{
idx = (vmm_frame_hint + i) % total_frames;
if(vmm_bit_get(idx) == (u64)0)
{
vmm_bit_set(idx, (u64)1);
vmm_frame_hint = idx + (u64)1;
return idx * (u64)4096;
}
i = i + (u64)1;
}
return (u64)0;
}
global void
vmm_free_frame(u64 addr)
{
vmm_bit_set(addr / (u64)4096, (u64)0);
}
global u64
vmm_frames_free(void)
{
u64 total_frames;
u64 i;
u64 n;
total_frames = (u64)0x40000000 / (u64)4096;
i = 0;
n = 0;
while(i < total_frames)
{
if(vmm_bit_get(i) == (u64)0)
{
n = n + (u64)1;
}
i = i + (u64)1;
}
return n;
}
```
- [ ] **Step 2: Write `vmm_init`**
```c
global void
vmm_init(void)
{
u64 i;
ptr entry;
u64 addr;
u64 size;
u32 type;
u64 start_frame;
u64 end_frame;
u64 f;
u64 reserved_end;
i = 0;
while(i < (u64)131072)
{
vmm_frame_bitmap[i] = (i8)0xff; /* start with everything marked
* used -- e820 usable ranges
* below clear only what's
* genuinely free */
i = i + (u64)1;
}
vmm_frame_hint = 0;
if(vmm_e820_count == (u64)0)
{
klog_write("kaboom: no usable memory map -- halting\n");
while(1)
{
halt();
}
}
reserved_end = heap_colosseum_limit(); /* 0x3000000 -- everything
* at or below this is
* already spoken for
* (kernel image, boot's own
* tables, the stack,
* heap_colosseum) and must
* never be handed out as a
* general frame */
i = 0;
while(i < vmm_e820_count)
{
entry = &vmm_e820_buf[0] + i * (u64)24;
addr = vmm_read_u64(entry, (u64)0);
size = vmm_read_u64(entry, (u64)8);
type = (u32)vmm_read_u32(entry, (u64)16);
if(type == (u32)1) /* usable RAM */
{
start_frame = (addr + (u64)4095) / (u64)4096;
end_frame = (addr + size) / (u64)4096; /* exclusive */
if(addr < reserved_end)
{
start_frame = reserved_end / (u64)4096;
}
if(end_frame > (u64)0x40000000 / (u64)4096)
{
end_frame = (u64)0x40000000 / (u64)4096; /* cap at 1gib */
}
f = start_frame;
while(f < end_frame)
{
vmm_bit_set(f, (u64)0);
f = f + (u64)1;
}
}
i = i + (u64)1;
}
}
```
- [ ] **Step 3: Wire into `kmain.nsc`**
Add `include "vmm.nsh";` to `kmain.nsc`'s include block, add `void vmm_init(void);`/`void vmm_e820_from_pvh(void);` forward decls, and change:
```c
alloc_init();
klog_write("kaboom: alloc ok\n");
paging_init();
klog_write("kaboom: paging ok\n");
```
to:
```c
alloc_init();
klog_write("kaboom: alloc ok\n");
paging_init();
klog_write("kaboom: paging ok\n");
vmm_e820_from_pvh();
vmm_init();
klog_write("kaboom: vmm ok\n");
```
(both `paging_init` and `vmm_init` run side by side for now — `paging.nsc` isn't deleted until Task 11, once `exec.nsc` no longer calls it).
- [ ] **Step 4: Build, boot, verify via `klogs`**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot, `qemu_wait_serial(text="$")`, `qemu_serial_send("klogs\n")`, confirm `kaboom: vmm ok` appears (proving `vmm_init` didn't hit the halt path).
- [ ] **Step 5: Commit**
```bash
git add src/kernel/vmm.nsc src/kernel/vmm.nsh src/kernel/kmain.nsc
git commit -F- <<'EOF'
ADD: vmm physical frame bitmap allocator
MOD: vmm_alloc_frame/vmm_free_frame hand out real 4KiB frames from a
bitmap seeded by the real pvh memory map, capped at the first 1GiB
(what boot.s's identity map covers) and reserving everything
alloc_init/heap_colosseum already claimed below 0x3000000. runs
alongside the still-live paging.nsc scratch pool for now -- retired
once exec.nsc no longer calls it (Task 11).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 4: `/int/vmm` diagnostic file
**Files:**
- Modify: `src/kernel/virtfs.nsc` (add `virtfs_gen_vmm`, one dispatch line)
**Interfaces:**
- Consumes: `vmm_frames_free()` (Task 3).
- Produces: a live, human-readable `cat /int/vmm` — the main tool every later task's own tests use to confirm frame counts before/after create/destroy.
This is small on its own merits, but every later task's regression tests depend on being able to observe the allocator's live state the way `/int/mem` already lets you observe `heap_colosseum`'s — building it now, once, pays for itself immediately.
- [ ] **Step 1: Write `virtfs_gen_vmm`, mirroring `virtfs_gen_mem`'s exact style**
Add right after the existing `virtfs_gen_mem` function (`src/kernel/virtfs.nsc`, currently ending around line 241):
```c
u64 vmm_frames_free(void);
i32
virtfs_gen_vmm(ptr buf)
{
u64 off;
off = 0;
off = virtfs_put_str(buf, off, "frames free: ");
off = virtfs_put_dec(buf, off, vmm_frames_free());
off = virtfs_put_str(buf, off, " (");
off = virtfs_put_dec(buf, off, vmm_frames_free() * (u64)4096);
off = virtfs_put_str(buf, off, " bytes)\n");
return (i32)off;
}
```
- [ ] **Step 2: Add the dispatch line**
In `virtfs_read`'s `kfs_int_dir_lba` branch (currently lines 346-365), add right after the existing `"cpu"` check:
```c
if(virtfs_name_eq(name, name_len, "vmm", (u64)3) == 1)
{
return virtfs_gen_vmm(buf);
}
```
- [ ] **Step 3: Build, boot, verify**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot, `qemu_wait_serial(text="$")`, `qemu_serial_send("cat /int/vmm\n")`, confirm it prints a sane, non-zero free-frame count and matching byte total.
- [ ] **Step 4: Commit**
```bash
git add src/kernel/virtfs.nsc
git commit -F- <<'EOF'
ADD: /int/vmm - live frame-allocator diagnostics
MOD: cat /int/vmm reports vmm's live free-frame count, the same
style /int/mem already reports heap_colosseum's usage -- the tool
every later vmm task's own regression tests use to confirm frames
actually get returned on process teardown.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 5: Page-table asm primitives + saved kernel CR3
**Files:**
- Create: `src/kernel/vmm_asm.s`
- Modify: `src/kernel/kmain.nsc` (save `kernel_cr3` at the very start)
- Modify: `mk/b.sh`
**Interfaces:**
- Produces: `u64 get_cr3(void);`, `void load_cr3(u64 pml4_phys);`, `ptr get_pd_table(void);` (same two lines `paging_asm.s` already has — moved here, not touched there yet, since `paging.nsc` still needs it until Task 11), `global u64 kernel_cr3;` (`kmain.nsc`) — Task 6 reads this to know what to restore CR3 to for the boot-time `sys_exec("sh", ...)` call (no parent process slot exists yet at that point).
- [ ] **Step 1: Write `vmm_asm.s`**
```asm
/*
* three tiny wrappers vmm.nsc needs -- nsc has no inline asm and no
* way to reference cr3 or a raw assembler label from nsc directly,
* only through a real extern'd function (same pattern paging_asm.s
* already used for get_pd_table/invlpg).
*/
.section .text
.global get_cr3
.type get_cr3, @function
get_cr3:
movq %cr3, %rax
ret
.global load_cr3
.type load_cr3, @function
load_cr3:
movq %rdi, %cr3
ret
.global get_pd_table
.type get_pd_table, @function
get_pd_table:
leaq pd(%rip), %rax
ret
```
- [ ] **Step 2: Save `kernel_cr3` at `kmain`'s very start**
In `kmain.nsc`, add `u64 get_cr3(void);` to the forward-decl block and `global u64 kernel_cr3;` near the top of the file (outside `kmain` itself, a real global — Task 6's `exec.nsc` reads it). As the very first line of `kmain`'s own body (before `klog_init()`):
```c
kernel_cr3 = get_cr3();
```
- [ ] **Step 3: Add to the build**
In `mk/b.sh`, add `"$as" $asflgs "$src/kernel/vmm_asm.s" -o "$out/vmm_asm.o"` right after the existing `paging_asm.o` line, and add `"$out/vmm_asm.o"` to the final `"$ld"` object list (right after `"$out/paging_asm.o"`).
- [ ] **Step 4: Build, boot, verify no regression**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot, confirm the shell prompt still comes up (`qemu_wait_serial(text="$")`) and `klogs` still shows the normal full boot sequence — this task adds no new observable behavior yet, only plumbing Task 6 needs.
- [ ] **Step 5: Commit**
```bash
git add src/kernel/vmm_asm.s src/kernel/kmain.nsc mk/b.sh
git commit -F- <<'EOF'
ADD: vmm_asm.s - cr3 read/write + pd-table accessor
MOD: get_cr3/load_cr3 let exec.nsc switch address spaces (Task 6);
kernel_cr3, saved once at kmain's very start, is what a boot-time
sys_exec("sh", ...) call restores to on return, since no parent
process slot exists yet at that depth.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 6: `vmm_create_addrspace` + `vmm_map_page`
**Files:**
- Modify: `src/kernel/vmm.nsc`, `src/kernel/vmm.nsh`
**Interfaces:**
- Consumes: `vmm_alloc_frame()` (Task 3), `get_pd_table()` (Task 5).
- Produces: `u64 vmm_create_addrspace(void);` (returns a new pml4 physical address, or `(u64)0` on failure with every partial frame already freed), `i32 vmm_map_page(u64 pml4_phys, u64 vaddr, u64 paddr);` (returns 1 on success, 0 if `vaddr` falls outside the private region `0x400000`-`0x1ffffff` or a page-table-structure frame couldn't be allocated) — Task 8 calls both directly.
This is the core of the whole project — re-splitting the existing 1GiB identity map (pd index 0-1 and 16-23 shared by value, 2-15 private per-process) exactly as the spec's architecture section lays out.
- [ ] **Step 1: Write the pd-index constants and the shared-copy logic**
```c
u64 get_cr3(void);
void load_cr3(u64 pml4_phys);
ptr get_pd_table(void);
global u64
vmm_create_addrspace(void)
{
u64 pml4_phys;
u64 pdpt_phys;
u64 pd_phys;
ptr boot_pd;
ptr new_pd;
u64 i;
u64 word;
pml4_phys = vmm_alloc_frame();
pdpt_phys = vmm_alloc_frame();
pd_phys = vmm_alloc_frame();
if(pml4_phys == (u64)0 || pdpt_phys == (u64)0 || pd_phys == (u64)0)
{
if(pml4_phys != (u64)0) { vmm_free_frame(pml4_phys); }
if(pdpt_phys != (u64)0) { vmm_free_frame(pdpt_phys); }
if(pd_phys != (u64)0) { vmm_free_frame(pd_phys); }
return (u64)0;
}
/* zero all three fresh frames -- every pml4/pdpt/pd entry starts
* not-present (bit 0 clear) until explicitly set below */
i = 0;
while(i < (u64)512)
{
*((ptr)pml4_phys + i * (u64)8) = 0;
*((ptr)pdpt_phys + i * (u64)8) = 0;
*((ptr)pd_phys + i * (u64)8) = 0;
i = i + (u64)1;
}
*(ptr)pml4_phys = (i64)(pdpt_phys | (u64)0x03); /* present+writable */
*(ptr)pdpt_phys = (i64)(pd_phys | (u64)0x03);
/* copy the shared indices (0-1: kernel low memory/stack; 16-23:
* heap_colosseum) BY VALUE from the boot-time pd -- these are
* still 2mib huge-page entries, identical physical==virtual
* mapping every process shares. indices 2-15 (the private region)
* stay zero/not-present, filled in lazily by vmm_map_page. */
boot_pd = get_pd_table();
new_pd = (ptr)pd_phys;
i = 0;
while(i < (u64)2)
{
word = (u64)*(boot_pd + i * (u64)8);
*(new_pd + i * (u64)8) = (i64)word;
i = i + (u64)1;
}
i = 16;
while(i < (u64)24)
{
word = (u64)*(boot_pd + i * (u64)8);
*(new_pd + i * (u64)8) = (i64)word;
i = i + (u64)1;
}
return pml4_phys;
}
```
- [ ] **Step 2: Write `vmm_map_page`**
```c
global i32
vmm_map_page(u64 pml4_phys, u64 vaddr, u64 paddr)
{
ptr pd;
u64 pd_index;
ptr pt_entry_in_pd;
u64 pt_phys;
ptr pt;
u64 pt_index;
u64 i;
if(vaddr < (u64)0x400000 || vaddr >= (u64)0x2000000)
{
return 0; /* outside the private region entirely */
}
/* pml4_phys's own entry 0 always points at this addrspace's one
* pdpt, whose entry 0 always points at this addrspace's one pd --
* both fixed by vmm_create_addrspace, never touched again, so
* walking straight to the pd needs no further pml4/pdpt indexing. */
pd = (ptr)((u64)*(ptr)((u64)*(ptr)pml4_phys & ~(u64)0xfff) & ~(u64)0xfff);
pd_index = vaddr / (u64)0x200000;
pt_entry_in_pd = pd + pd_index * (u64)8;
if(((u64)*pt_entry_in_pd & (u64)1) == (u64)0)
{
/* this 2mib slice has never been touched by this process
* before -- allocate a fresh per-process pt for it */
pt_phys = vmm_alloc_frame();
if(pt_phys == (u64)0)
{
return 0;
}
i = 0;
while(i < (u64)512)
{
*((ptr)pt_phys + i * (u64)8) = 0;
i = i + (u64)1;
}
*pt_entry_in_pd = (i64)(pt_phys | (u64)0x03);
}
pt = (ptr)((u64)*pt_entry_in_pd & ~(u64)0xfff);
pt_index = (vaddr / (u64)4096) % (u64)512;
*(pt + pt_index * (u64)8) = (i64)(paddr | (u64)0x03);
return 1;
}
```
- [ ] **Step 2b: Declare both in `vmm.nsh`**
```c
u64 vmm_create_addrspace(void);
i32 vmm_map_page(u64 pml4_phys, u64 vaddr, u64 paddr);
```
- [ ] **Step 3: Verify via a temporary `kmain.nsc` smoke test**
After `vmm_init()`'s call in `kmain.nsc` (temporary, removed at the end of this step):
```c
u64 as1;
u64 f1;
ptr w;
as1 = vmm_create_addrspace();
klog_write(as1 != (u64)0 ? "kaboom: addrspace create ok\n" : "kaboom: addrspace create FAILED\n");
f1 = vmm_alloc_frame();
if(vmm_map_page(as1, (u64)0x400000, f1) == 1)
{
klog_write("kaboom: map ok\n");
w = (ptr)f1; /* write through the kernel's own identity alias of
the same physical frame, since cr3 hasn't switched
yet */
*w = (i64)0x1122334455667788;
klog_write("kaboom: wrote via identity alias\n");
}
else
{
klog_write("kaboom: map FAILED\n");
}
```
Build, boot, `klogs`, confirm all three "ok" lines appear (proving frame allocation, address-space creation, and the private-region mapping all work end to end). Remove this temporary block once confirmed (Task 8 replaces it with the real integration).
- [ ] **Step 4: Commit**
```bash
git add src/kernel/vmm.nsc src/kernel/vmm.nsh
git commit -F- <<'EOF'
ADD: vmm_create_addrspace/vmm_map_page - per-process page tables
MOD: every new address space gets a fresh pml4/pdpt/pd, sharing the
boot-time pd's own low-memory and heap_colosseum entries by value and
mapping its own 0x400000-0x1ffffff private region on demand, one
per-process pt per touched 2mib slice.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 7: `vmm_destroy_addrspace`
**Files:**
- Modify: `src/kernel/vmm.nsc`, `src/kernel/vmm.nsh`
**Interfaces:**
- Consumes: `vmm_free_frame()` (Task 3).
- Produces: `void vmm_destroy_addrspace(u64 pml4_phys);` — Task 8 calls this once `elf_call_entry` returns.
- [ ] **Step 1: Write `vmm_destroy_addrspace`**
```c
global void
vmm_destroy_addrspace(u64 pml4_phys)
{
ptr pd;
u64 pdpt_phys;
u64 i;
u64 pde;
u64 pt_phys;
ptr pt;
u64 j;
u64 pte;
pdpt_phys = (u64)*(ptr)pml4_phys & ~(u64)0xfff;
pd = (ptr)((u64)*(ptr)pdpt_phys & ~(u64)0xfff);
/* free only the private indices (2-15) and whatever per-process
* pt/frames they point at -- indices 0-1/16-23 are the shared
* kernel mapping, never owned by this process, never freed here */
i = 2;
while(i < (u64)16)
{
pde = (u64)*(pd + i * (u64)8);
if((pde & (u64)1) == (u64)1)
{
pt_phys = pde & ~(u64)0xfff;
pt = (ptr)pt_phys;
j = 0;
while(j < (u64)512)
{
pte = (u64)*(pt + j * (u64)8);
if((pte & (u64)1) == (u64)1)
{
vmm_free_frame(pte & ~(u64)0xfff);
}
j = j + (u64)1;
}
vmm_free_frame(pt_phys);
}
i = i + (u64)1;
}
vmm_free_frame((u64)pd);
vmm_free_frame(pdpt_phys);
vmm_free_frame(pml4_phys);
}
```
- [ ] **Step 1b: Declare in `vmm.nsh`**
```c
void vmm_destroy_addrspace(u64 pml4_phys);
```
- [ ] **Step 2: Verify via a temporary `kmain.nsc` extension of Task 6's smoke test**
Extend Task 6's (already-removed) smoke test temporarily: create an address space, map a page, note `vmm_frames_free()` before creating and after destroying, confirm they match:
```c
u64 free_before;
u64 as2;
free_before = vmm_frames_free();
as2 = vmm_create_addrspace();
vmm_map_page(as2, (u64)0x400000, vmm_alloc_frame());
vmm_destroy_addrspace(as2);
klog_write(vmm_frames_free() == free_before ? "kaboom: destroy ok\n" : "kaboom: destroy FAILED (leak)\n");
```
Build, boot, `klogs`, confirm `kaboom: destroy ok`. Remove this temporary block (Task 8 exercises the real path).
- [ ] **Step 3: Commit**
```bash
git add src/kernel/vmm.nsc src/kernel/vmm.nsh
git commit -F- <<'EOF'
ADD: vmm_destroy_addrspace - full per-process teardown
MOD: frees every frame a process's own private pd entries (2-15)
point at, plus its pt/pd/pdpt/pml4 structure frames themselves --
never touches the shared indices (0-1, 16-23). live-verified: create+
map+destroy returns vmm_frames_free() to exactly its pre-create value.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 8: `exec.nsc` integration — the real `sys_exec` rewrite
**Files:**
- Modify: `src/kernel/exec.nsc:280-302` (delete `exec_shwin_live`/`exec_shwin_restore`), `src/kernel/exec.nsc:316-560` (`sys_exec` body), `src/kernel/exec.nsh` (add `kaboom_proc_pml4[10]`)
- Modify: `src/kernel/kmain.nsc` (remove the now-unused Task 6/7 smoke-test remnants if any linger; confirm none do)
**Interfaces:**
- Consumes: `vmm_create_addrspace`/`vmm_map_page`/`vmm_destroy_addrspace` (Tasks 6-7), `load_cr3`/`get_cr3` (Task 5), `kernel_cr3` (Task 5, `kmain.nsc`).
- Produces: the real, working per-process `sys_exec` — every later task's regression tests exercise this directly.
This is the highest-risk task in the plan: it replaces the *only* path anything executes through. Do not skip Step 4's full regression pass.
- [ ] **Step 1: Extend the process table**
In `exec.nsh`, add one more parallel array alongside the existing `kaboom_proc_pid`/`kaboom_proc_name`:
```c
global u64 kaboom_proc_pml4[10];
```
- [ ] **Step 2: Delete the retired shwin machinery**
In `exec.nsc`, delete lines 280-302 (`exec_shwin_live`'s declaration and `exec_shwin_restore`) in full. Delete the `include "paging.nsh";` line (currently line 86) — `exec.nsc` no longer calls anything from `paging.nsc`. Add `include "vmm.nsh";` in its place.
- [ ] **Step 3: Rewrite `sys_exec`'s body**
Replace the whole function (currently lines 316-560) with:
```c
global i64
sys_exec(ptr path, u64 path_len, i64 argc, ptr argv)
{
i32 n;
u64 sz;
u64 entry;
i32 from_depth;
i64 result;
i8 interp_path[64];
u64 interp_len;
ptr new_argv[16];
i64 new_argc;
i32 ai;
i8 resolved_path[80];
ptr use_path;
u64 use_path_len;
i32 joined_len;
u64 need;
u64 zi;
ptr nb;
u32 fd_keep;
u64 new_pml4;
u64 saved_cr3;
u64 seg_i;
u64 ph_off_local;
u64 e_phoff;
u16 e_phentsize;
u16 e_phnum;
u32 p_type;
u64 p_offset;
u64 p_vaddr;
u64 p_filesz;
u64 p_memsz;
u64 page_va;
u64 page_pa;
u64 copy_off;
u64 word;
u64 bi;
u8 byteval;
/* path resolution, permission check, and reading the file into nb
* -- unchanged from the existing code. */
use_path = path;
use_path_len = path_len;
n = kfs_resolve(path, path_len);
if(n < 0)
{
n = kfs_dir_find_in(kfs_bin_dir_lba, path, path_len);
if(n >= 0)
{
joined_len = kfs_join_path(&resolved_path[0], (u64)80, "/bin", (u64)4, path, path_len);
if(joined_len >= 0)
{
use_path = &resolved_path[0];
use_path_len = (u64)joined_len;
}
}
}
if(n < 0)
{
return -1;
}
if(kfs_check_perm(n, (u64)4) == 0)
{
return -1;
}
need = kfs_file_size(n);
if(need > (u64)8450048)
{
return -1;
}
need = (need + (u64)7) & ~(u64)7;
if(need < (u64)63488)
{
need = (u64)63488;
}
if(need <= (u64)65536)
{
nb = exec_filebuf;
}
else
{
nb = arena_alloc(&fs_colosseum, need);
if(nb == (ptr)0)
{
return -1;
}
}
zi = 0;
while(zi < need)
{
*(nb + zi) = 0;
zi = zi + (u64)8;
}
sz = kfs_read_file(n, nb);
if(sz == (u64)0)
{
return -1;
}
if(sz >= (u64)2 && ptr_byte_at(nb, (u64)0) == (u8)35 && ptr_byte_at(nb, (u64)1) == (u8)33)
{
/* shebang handling -- unchanged from the existing code, except
* the comment below (the old one referenced "sh already live
* in its own window", which no longer exists). */
interp_len = shebang_parse_interp(nb, sz, &interp_path[0]);
if(interp_len == (u64)0)
{
return -1;
}
new_argv[0] = &interp_path[0];
new_argv[1] = use_path;
new_argc = 2;
ai = 1;
while(ai < (i32)argc && new_argc < (i64)16)
{
new_argv[new_argc] = exec_argv_get(argv, ai);
new_argc = new_argc + (i64)1;
ai = ai + 1;
}
/* the interpreter here is, in every real case, sh -- exec'd
* recursively below, which now gets its own real address space
* like any other process, no special case. */
if(exec_shebang_depth >= 8)
{
return -1;
}
exec_shebang_depth = exec_shebang_depth + 1;
result = sys_exec(&interp_path[0], interp_len, new_argc, (ptr)&new_argv[0]);
exec_shebang_depth = exec_shebang_depth - 1;
return result;
}
/* everything from here down replaces the old window/shwin logic
* (previously lines 495-531) entirely. */
if(elf_validate(nb) == 0 || sz < (u64)64)
{
return -1;
}
e_phoff = elf_read_u64_pub(nb, (u64)32); /* see Step 3b below --
elf.nsc needs 3 new
exports for this */
e_phentsize = elf_read_u16_pub(nb, (u64)54);
e_phnum = elf_read_u16_pub(nb, (u64)56);
if(elf_segments_ok(nb, sz, elf_read_u64_pub(nb, (u64)24), e_phoff, e_phentsize, e_phnum) == 0)
{
return -1;
}
new_pml4 = vmm_create_addrspace();
if(new_pml4 == (u64)0)
{
return -1;
}
/* map + copy every PT_LOAD segment, one 4kib page at a time,
* rolling back the whole address space on any allocation failure
* partway through */
seg_i = 0;
while(seg_i < (u64)e_phnum)
{
ph_off_local = e_phoff + seg_i * (u64)e_phentsize;
p_type = elf_read_u32_pub(nb, ph_off_local);
if(p_type == (u32)1)
{
p_offset = elf_read_u64_pub(nb, ph_off_local + (u64)8);
p_vaddr = elf_read_u64_pub(nb, ph_off_local + (u64)16);
p_filesz = elf_read_u64_pub(nb, ph_off_local + (u64)32);
p_memsz = elf_read_u64_pub(nb, ph_off_local + (u64)40);
page_va = p_vaddr & ~(u64)0xfff;
while(page_va < p_vaddr + p_memsz)
{
page_pa = vmm_alloc_frame();
if(page_pa == (u64)0 || vmm_map_page(new_pml4, page_va, page_pa) == 0)
{
vmm_destroy_addrspace(new_pml4);
return -1;
}
/* zero the whole frame first (a fresh frame may still
* hold a PREVIOUS process's data -- see this plan's
* own Review Focus on this exact point), then copy
* whatever file bytes this page actually covers */
bi = 0;
while(bi < (u64)4096)
{
*((ptr)page_pa + bi) = 0;
bi = bi + (u64)8;
}
copy_off = 0;
while(copy_off < (u64)4096)
{
if(page_va + copy_off >= p_vaddr && page_va + copy_off < p_vaddr + p_filesz)
{
word = elf_read_u64_pub(nb, p_offset + (page_va + copy_off - p_vaddr));
*((ptr)page_pa + copy_off) = (i64)word;
}
copy_off = copy_off + (u64)8;
}
page_va = page_va + (u64)4096;
}
}
seg_i = seg_i + (u64)1;
}
entry = elf_read_u64_pub(nb, (u64)24);
from_depth = proc_push(use_path, use_path_len);
kaboom_proc_pml4[from_depth] = new_pml4;
saved_cr3 = (from_depth == 0) ? kernel_cr3 : kaboom_proc_pml4[from_depth - 1];
load_cr3(new_pml4);
fd_keep = fd_open_mask();
result = elf_call_entry((ptr)entry, argc, argv);
fd_close_unless(fd_keep);
load_cr3(saved_cr3);
vmm_destroy_addrspace(new_pml4);
proc_pop(from_depth);
return result;
}
```
**Note for the implementer:** the `_pub`-suffixed `elf_read_u64`/`elf_read_u32`/`elf_read_u16`/`elf_validate` calls above don't exist yet — Task 9 exports them from `elf.nsc` (currently `static`-equivalent/file-local helpers there). Do Task 9 either just before this step or as part of this same task if working sequentially — the two are tightly coupled and this task's own code won't compile without Task 9's exports. `u64 kernel_cr3;` needs an `extern`-style forward reference at the top of `exec.nsc` (`void` isn't right for a global — just reference it directly the way `exec.nsc` already references `alloc.nsh`'s `heap_colosseum`-adjacent globals: a plain `global u64 kernel_cr3;` re-declaration is wrong here since it's defined in `kmain.nsc`; instead declare it the way this codebase already threads a global defined in one `.nsc` file to another — check `kaboom_next_pid`'s own cross-file declaration pattern in `exec.nsh` for the exact established idiom and mirror it for `kernel_cr3`, moving its definition from `kmain.nsc` into `exec.nsh`/`exec.nsc` if that idiom requires the definition and declaration to travel together).
- [ ] **Step 4: Full regression pass — build, boot, exercise everything**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot and verify, in order:
1. Shell prompt comes up (`qemu_wait_serial(text="$")`).
2. `ls`, `cat /proc/version`, `pwd` — ordinary commands still work.
3. `cat /int/vmm` — note the free-frame count.
4. `sh` (nested), `cat /proc/ps` (from inside it) — confirm both levels show correctly, `exit` back out.
5. `cat /int/vmm` again — confirm the free-frame count matches step 3's (the nested `sh` and its own `cat` both tore down cleanly).
6. Nest `sh` 8 times in a row (matching this session's own earlier-established test of the old `exec_shwin` limit) and confirm it still succeeds at depth matching whatever the *new* frame-pool-driven ceiling turns out to be (record the actual number reached — this is expected to differ from the old hardcoded 8, since the limit is now real frame exhaustion, not a counter; document whatever real number is observed in the commit message).
7. A shebang chain 8 levels deep (same test this session already proved once) — confirm it still works.
- [ ] **Step 5: Commit**
```bash
git add src/kernel/exec.nsc src/kernel/exec.nsh src/kernel/kmain.nsc
git commit -F- <<'EOF'
MOD: exec.nsc - real per-process address spaces replace the shwin trick
MOD: sys_exec now builds a real address space per process (vmm.nsc),
switching cr3 around elf_call_entry instead of remapping a shared
physical window. kaboom_proc_pml4[10] extends the existing process
table. full regression pass: ordinary commands, nested sh + /proc/ps,
[N]-deep sh nesting, 8-deep shebang chains all confirmed live.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 9: `elf.nsc` — export byte-readers, retire the window check
**Files:**
- Modify: `src/kernel/elf.nsc`
**Interfaces:**
- Produces: `global u64 elf_read_u64_pub(ptr buf, u64 off);` etc. (Task 8 depends on these) — or, simpler and preferred: just mark the *existing* `elf_read_u16`/`u32`/`u64`/`elf_validate` functions `global` in place (they're plain functions already, `global` is the only change needed — rename references in Task 8's code above from the `_pub` suffix back to the real names `elf_read_u64`/`elf_read_u32`/`elf_read_u16`/`elf_validate` once this task lands; the `_pub` suffix in Task 8's draft above exists only to flag "this doesn't exist yet," not as a real naming convention to keep).
- [ ] **Step 1: Mark the byte-readers `global`**
In `elf.nsc`, change `u16 elf_read_u16(ptr buf, u64 off)` (line 28) to `global u16`, `u32 elf_read_u32` (line 38) to `global u32`, `u64 elf_read_u64` (line 53) to `global u64`. `elf_validate` (line 72) is already `global`.
- [ ] **Step 2: Delete `elf_window_lo`/`elf_window_hi`/`elf_window`**
Delete `elf_window_lo` (lines 92-98), `elf_window_hi` (lines 100-105), and `elf_window` (lines 202-210) in full — no caller needs them once `sys_exec` no longer branches on which fixed window an entry point falls in.
- [ ] **Step 3: Replace `elf_segments_ok`'s window check with a private-region bounds check**
In `elf_segments_ok` (currently lines 121-189), replace:
```c
lo = elf_window_lo(e_entry);
if(lo == (u64)0)
{
return 0;
}
hi = elf_window_hi(lo);
```
with:
```c
lo = (u64)0x400000;
hi = (u64)0x2000000;
if(e_entry < lo || e_entry >= hi)
{
return 0;
}
```
(the rest of the function — the phdr-table bounds check, the per-segment `p_offset`/`p_filesz`/`p_vaddr`/`p_memsz` checks against `lo`/`hi` — is already written generically against `lo`/`hi` and needs no further change; only the two now-unused local variable declarations `lo`/`hi`'s *source* changes, not their later use).
- [ ] **Step 4: Delete `elf_load` — Task 8 replaced its caller**
`elf_load` (lines 212-308) copied segments directly to a fixed physical address — `sys_exec` (Task 8) now does its own per-page copy through `vmm_map_page`, so nothing calls `elf_load` any more. Delete it in full. (If Task 8 was implemented expecting to call `elf_load`, stop and reconcile: the plan's own Task 8 draft does the copy inline via `vmm_map_page` + a manual copy loop specifically so it can allocate one physical frame at a time rather than needing one large contiguous buffer the way `elf_load` assumed — this deletion is intentional, not an oversight.)
- [ ] **Step 5: Build (compile-only check, since `elf.nsc`'s callers were already updated in Task 8)**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh
```
Confirm it compiles clean (Task 8's own build/boot verification already covers full behavior — this step exists only to isolate an `elf.nsc`-specific compile error from a `sys_exec`-specific one if something goes wrong).
- [ ] **Step 6: Commit**
```bash
git add src/kernel/elf.nsc
git commit -F- <<'EOF'
MOD: elf.nsc - drop the two-window model, export byte-readers
MOD: elf_window_lo/elf_window_hi/elf_window/elf_load are all deleted
-- there's only one private region now (0x400000-0x1ffffff), checked
directly in elf_segments_ok, and sys_exec (exec.nsc) does its own
per-page copy through vmm_map_page instead of elf_load's single
fixed-address copy.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 10: Unify the link scripts
**Files:**
- Create: `src/user/user.ld`
- Delete: `src/user/user_shell.ld`, `src/user/user_prog.ld`
- Modify: `mk/bu.sh`
**Interfaces:**
- Consumes: nothing from earlier tasks directly — independent of the vmm work itself, but only *safe* to land once Task 8 is live (before that, `sh` and ordinary programs still need different addresses).
- [ ] **Step 1: Write `user.ld`**
```
/*
* link script for every userspace program, sh included -- now that
* each process gets its own real address space (vmm.nsc), isolation
* comes from cr3, not address partitioning, so there's no reason for
* sh to need a different fixed address from anything it runs any
* more (see the old user_shell.ld/user_prog.ld's own history for why
* that split existed in the first place).
*
* two PT_LOAD segments (R+X text/rodata, R+W bss), tiny linker page
* size -- unchanged reasoning from the old scripts: avoids ld's "RWX
* permissions" warning at near-zero cost, since this loader has no
* real use for page alignment.
*/
PHDRS
{
rx PT_LOAD FLAGS(5); /* R+X, never W */
rw PT_LOAD FLAGS(6); /* R+W, never X */
}
ENTRY(_start)
SECTIONS
{
. = 0x400000;
.text : { *(.text) } :rx
.rodata : { *(.rodata) } :rx
. = ALIGN(8);
.bss : { *(.bss) *(COMMON) } :rw
/DISCARD/ : { *(.note.*) *(.comment) *(.eh_frame) }
}
```
- [ ] **Step 2: Update `mk/bu.sh`**
Replace:
```sh
if [ "$name" = "sh" ]; then
ldscript="$userdir/user_shell.ld"
else
ldscript="$userdir/user_prog.ld"
fi
```
with:
```sh
ldscript="$userdir/user.ld"
```
Update the file's own top comment (currently: "sh links against user_shell.ld (0x400000); everything else links against user_prog.ld (0x600000)") to: "every program links against user.ld (0x400000) -- now that each process gets its own real address space (vmm.nsc), there's no reason for sh to need a different address from anything it runs."
- [ ] **Step 3: Delete the old scripts**
```bash
git rm src/user/user_shell.ld src/user/user_prog.ld
```
- [ ] **Step 4: Build, boot, verify**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot, confirm `sh` and every coreutil (`ls`, `cat`, `cp`, `chmod`, etc.) still run correctly — same regression checklist as Task 8 Step 4, abbreviated (this task shouldn't change behavior at all if Task 8 landed correctly, only where each binary links).
- [ ] **Step 5: Commit**
```bash
git add -A
git commit -F- <<'EOF'
MOD: unify user_shell.ld/user_prog.ld into one user.ld
MOD: every userspace program (sh included) now links at 0x400000 --
real per-process address spaces (vmm.nsc) made the old two-address
split unnecessary; isolation comes from cr3, not address partitioning.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 11: Retire `paging.nsc`
**Files:**
- Delete: `src/kernel/paging.nsc`, `src/kernel/paging.nsh`, `src/kernel/paging_asm.s`
- Modify: `mk/b.sh`, `src/kernel/kmain.nsc`
**Interfaces:**
- Consumes: nothing — this is pure cleanup, safe only once Task 8 already removed every call into `paging.nsc` from `exec.nsc`.
- [ ] **Step 1: Confirm nothing still references it**
```bash
cd /root/work/kaboom
grep -rn "paging_init\|paging_alloc_frame\|paging_free_frame\|page_current_phys\|page_remap_2mb\|paging\.nsh" src/ mk/
```
Expect zero hits outside `paging.nsc`/`paging.nsh`/`paging_asm.s` themselves and `mk/b.sh`'s own build lines for them. If anything else still references these, stop — Task 8 missed something, go fix it there first, don't patch around it here.
- [ ] **Step 2: Remove `paging_init()`'s call from `kmain.nsc`**
Delete the `paging_init();`/`klog_write("kaboom: paging ok\n");` lines, and the `void paging_init(void);` forward decl, from `kmain.nsc`.
- [ ] **Step 3: Delete the files and their build-list entries**
```bash
git rm src/kernel/paging.nsc src/kernel/paging.nsh src/kernel/paging_asm.s
```
In `mk/b.sh`, remove the `paging.nsc`/`paging.s`/`paging.o` compile+assemble lines, the `paging_asm.s`/`paging_asm.o` assemble line, and both `.o` files from the final `"$ld"` object list.
- [ ] **Step 4: Build, boot, verify**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot, confirm `klogs` shows a clean boot sequence (no more `kaboom: paging ok` line — that's expected and correct now), and a quick smoke test (`ls`, `cat /int/vmm`) still works.
- [ ] **Step 5: Commit**
```bash
git add -A
git commit -F- <<'EOF'
DEL: retire paging.nsc - superseded by vmm.nsc's real address spaces
MOD: the scratch-frame-pool/exec_shwin remap trick existed only to
let a nested sh reuse one shared window without colliding with the
copy it's nested inside -- vmm.nsc's per-process address spaces make
that unnecessary for any process, sh included, with no special case.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 12: Full regression + isolation test pass (spec's testing plan)
**Files:** none changed — this task is pure verification, matching the spec's own "Testing Plan" section point for point.
**Interfaces:** consumes the fully-integrated system from Tasks 1-11.
- [ ] **Step 1: Every existing coreutil**
Boot, run `ls`, `cat`, `cp`, `mv`, `mkdir`, `rmdir`, `touch`, `rm`, `chmod`, `perms`, `pwd`, `date`, `info`, `klogs`, `ps`, `4c` (any small script) — confirm each behaves exactly as before this project (no behavior change expected, only the underlying execution mechanism changed).
- [ ] **Step 2: Deep `sh` nesting + `/proc/ps` correctness**
Nest `sh` repeatedly until it fails, recording the real depth reached (this replaces the old hardcoded 8-frame ceiling with whatever the frame allocator's real capacity now allows — expect it to be larger, since frames are 4KiB now instead of 2MiB). At the deepest level, `cat /proc/ps` and confirm every level shows correctly (same test this session already used to verify the process-table fix earlier).
- [ ] **Step 3: Shebang chain depth**
Build an 8-level shebang chain (same construction this session already used once: `/s0` -> `/s1` -> ... -> `/s7` -> `/bin/echo`), confirm it still succeeds, then extend to 9 levels and confirm it still fails cleanly with the shell remaining responsive afterward.
- [ ] **Step 4: Two-sequential-exec isolation proof**
Write two tiny test programs (or reuse `4c` scripts) that each write a distinct, recognizable value to the same heap address and read it back — run them back to back, confirm the second never observes the first's value (proves real isolation, not an accidental leftover mapping; this is the concrete test for this plan's own Review Focus item #3 on frame zeroing).
- [ ] **Step 5: Frame-exhaustion rollback proof**
Temporarily shrink the usable-RAM cap in `vmm_init` (e.g. to a few hundred KB) to force real exhaustion, confirm a `sys_exec` mid-construction fails cleanly with `/int/vmm`'s free count unchanged from before the attempt, then revert the temporary cap and confirm normal execs succeed again. This directly exercises this plan's own Review Focus item #1.
- [ ] **Step 6: Boot-time `sys_exec("sh", ...)` restart proof**
`exit` the top-level shell from within `qemu_serial_send`, confirm `kmain`'s own respawn loop (`kaboom: exec sh failed` should NOT appear) brings up a fresh `sh` correctly, proving `kernel_cr3` restoration at depth 0 (this plan's own Review Focus item #4) works across more than one boot-time `sys_exec` call.
- [ ] **Step 7: Commit** (only if any fix was needed during this pass; otherwise this task produces no diff)
```bash
git add -A
git commit -F- <<'EOF'
MOD: fix issues found during full vmm regression pass
[describe whatever was actually found and fixed here]
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 13: Real-BIOS e820 path (`dynamite.s`)
**Files:**
- Modify: `src/boot/dynamite.s`
**Interfaces:**
- Consumes: `vmm_e820_add` (Task 2) — this task's whole job is calling it from the real-mode BIOS path instead of `vmm_e820_from_pvh`.
- Produces: `vmm_e820_buf`/`vmm_e820_count` populated identically to the PVH path, so `vmm_init` (Task 3) needs no changes at all to work with either source.
This task is independent of every other and deliberately ordered last — it only matters for the real-BIOS boot path (`mk/rb.sh`), which this session has never exercised directly, so it carries real, separate risk from the qemu `-kernel`/PVH path everything else in this plan was verified against.
- [ ] **Step 1: Add the real-mode INT 15h E820h loop to `dynamite.s`**
Before the existing protected-mode transition code, add (real-mode, 16-bit):
```asm
/* real bios e820 memory map -- same (addr, size, type) 24-byte
* entry shape vmm_e820_add (vmm.nsc) already expects, written
* directly into a fixed low-memory scratch buffer kmain reads
* after the mode transition (same "compute it once in real mode,
* hand it forward" pattern this file already uses for
* KERNEL_ENTRY_ADDR/KERNEL_LOAD_ADDR). */
xorl %ebx, %ebx
movl $e820_buf, %edi
movl $0, e820_count
e820_loop:
movl $0x0534d4150, %edx /* 'SMAP' */
movl $0xe820, %eax
movl $24, %ecx
int $0x15
jc e820_done /* carry set: error or end of list */
cmpl $0x0534d4150, %eax
jne e820_done /* bios didn't return 'SMAP': stop */
incl e820_count
addl $24, %edi
cmpl $256, e820_count
jae e820_done /* hit the same 256-entry ceiling
* vmm_e820_add enforces -- stop
* asking rather than overflow
* e820_buf */
testl %ebx, %ebx
jnz e820_loop /* ebx==0 means that was the last entry */
e820_done:
```
Add `.bss`/`.data` labels `e820_buf: .skip 6144` and `e820_count: .long 0` near dynamite's other scratch storage.
- [ ] **Step 2: Hand the buffer forward to `kmain`, same pattern as `KERNEL_ENTRY_ADDR`**
Check how `dynamite.s` currently hands `KERNEL_ENTRY_ADDR`/`KERNEL_LOAD_ADDR` forward into the relocated kernel (likely a fixed low-memory location the post-relocation kernel reads) and thread `e820_buf`'s physical address + `e820_count` through the same mechanism. Add a small `kmain.nsc`-side check: if booted via the BIOS path (some existing flag or detection this file already has for "which boot path got us here" — check `kmain.nsc`/`boot.s` for one; if none exists, this needs its own small addition here), call a new `vmm_e820_from_bios(u64 buf_addr, u64 count)` (mirrors `vmm_e820_from_pvh`, just copying `count` real BIOS-format entries — note the BIOS E820 entry is the *same* 24-byte `(addr, size, type, reserved)` shape already assumed, so this is closer to a straight copy than a re-parse) instead of `vmm_e820_from_pvh()`.
- [ ] **Step 3: Test via the real-BIOS boot path**
```bash
export PATH="/root/work/nscc/out:$PATH"
cd /root/work/kaboom
sh mk/b.sh && sh mk/bu.sh && sh mk/bl.sh && sh mk/bd.sh
```
Boot via `mk/rb.sh`'s own mechanism (`-drive` only, no `-kernel` — check that script's exact qemu invocation and mirror it through the `qemu_boot` MCP tool's `disk`/`extra_args` parameters, since this is a different boot path from every other task's verification in this plan), confirm the machine boots to a shell prompt and `cat /int/vmm` reports a sane free-frame count, same as the PVH path.
- [ ] **Step 4: Commit**
```bash
git add src/boot/dynamite.s
git commit -F- <<'EOF'
ADD: dynamite.s - real bios int 15h e820 memory detection
MOD: the real-bios boot path now detects actual installed ram via a
real-mode int 15h/e820h loop, writing into the same 24-byte entry
format vmm_e820_add already expects -- vmm_init needs no changes at
all to work with either boot path.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
```
---
### Task 14: Documentation sync
**Files:**
- Modify: `docs/kernel.btft` (the paging section and the "no pcb and no scheduler"/window-model language), `docs/build.btft` (if it mentions `user_shell.ld`/`user_prog.ld` by name — check first)
**Interfaces:** none — pure documentation, done last once every prior task is live-verified.
- [ ] **Step 1: Search for stale references**
```bash
cd /root/work/kaboom
grep -n "user_shell\.ld\|user_prog\.ld\|paging\.nsc\|exec_shwin\|0x600000\|two.window\|shared window" docs/*.btft
```
- [ ] **Step 2: Rewrite each stale section**
For `docs/kernel.btft`'s existing paging section (the one covering `paging.nsc`'s scratch pool / `exec_shwin`, and the "no pcb and no scheduler" process-table paragraph that still describes the retired model): replace with an accurate description of `vmm.nsc`'s real per-process address spaces — reuse this plan's own Architecture/spec language, scaled to `.btft`'s existing prose style (see this session's own earlier doc edits to `docs/kernel.btft`'s process-table paragraph for the established tone/tag conventions: `[tt]`/`[i]`/`[see name="..."]`).
For any `docs/build.btft` mention of `user_shell.ld`/`user_prog.ld`: update to `user.ld`, same one-link-script-for-everything language as Task 10's own commit message.
- [ ] **Step 3: Verify tag balance manually** (no `btf2html` in this sandbox — same manual check this session already used)
```bash
cd /root/work/kaboom
python3 -c "
import re
for fname in ['docs/kernel.btft', 'docs/build.btft']:
text = open(fname).read()
opens = len(re.findall(r'\[(tt|i|b)\]', text))
sees = len(re.findall(r'\[see name=\"[^\"]*\"\]', text))
closes = text.count('[e]')
print(fname, opens+sees, closes, 'OK' if opens+sees <= closes else 'CHECK')
"
```
- [ ] **Step 4: Commit and push**
```bash
git add docs/kernel.btft docs/build.btft
git commit -F- <<'EOF'
MOD: docs - describe vmm's real per-process address spaces
MOD: kernel.btft's paging/process-table sections and build.btft's
user.ld mention updated to match the retired two-window model.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EOF
GIT_SSH_COMMAND="ssh -i /root/.ssh/druid_deploy -o IdentitiesOnly=yes" git push origin trunk
```