do not edit — generated by btf.
git.druid.rocksindexdruid520kaboomdocs/superpowers/specs/2026-09-30-vmm-per-process-addrspace-design.md

docs/superpowers/specs/2026-09-30-vmm-per-process-addrspace-design.md


# per-process address spaces (vmm)
 
## context / motivation
 
kaboom has no scheduler and no per-process memory isolation. every
"process" is an ordinary called function running in the single shared
address space `boot.s` built at boot: one identity-mapped 1GiB region
(`pml4[0] -> pdpt[0] -> pd[0..511]`, each pd entry a 2MiB huge page,
`pd[i]` covering physical/virtual `i*0x200000` .. `i*0x200000+0x1fffff`
-- identity, since phys==virt everywhere in this map). within that,
exactly two fixed windows exist for userland code: `0x400000` (`sh`
only, `user_shell.ld`) and `0x600000`-`0x2000000` (everything else,
`user_prog.ld`), each with **one** live occupant at a time --
`paging.nsc`'s scratch-frame pool + `exec_shwin` remap trick exists
solely to let a *nested* `sh` reuse the first window without colliding
with the outer one it's nested inside.
 
this is the direct blocker for real daemons: a daemon sitting in the
"everything else" window would prevent any other program from running
at all while it's alive. this project replaces the two-window scheme
with real, private, per-process address spaces, so an arbitrary number
of processes (bounded only by physical memory) can be resident at
once. it is a prerequisite for the scheduler + daemon project that
follows it -- that project adds concurrency (a timer, TCBs, preemption,
the daemon mechanism); this one only adds isolation. execution stays
exactly as synchronous/blocking as it is today (`sys_exec` still calls
`elf_call_entry` and doesn't return until the child does).
 
decisions already made with the user, carried into this spec:
- **ring0-only.** no TSS, no user/kernel privilege split. isolation
  here means "processes can't collide with each other's memory by
  accident," not "processes can't attack each other" -- ring3 is a
  separate, later, well-scoped project (`gdt.nsc`'s own comment already
  anticipates it).
- **real e820/memory-map detection**, not a hardcoded RAM size --
  needed so the frame allocator never hands out a frame past the end
  of real, installed memory (exactly the bug class `paging.nsc`'s own
  history section already lived through once with the old scratch
  pool).
- **4KiB pages for process-private memory** (not 2MiB huge pages) --
  a 16KB program shouldn't burn a full 2MiB frame; on a 64-128MB
  machine that's the difference between a handful of resident
  processes and dozens.
 
## goals
 
- every `sys_exec`'d process gets its own private virtual region,
  backed by physical frames nothing else can touch, allocated on
  demand and fully reclaimed when the process exits.
- a general-purpose physical frame allocator (alloc *and* free, in any
  order) replacing the old LIFO-only scratch pool.
- real physical memory detection (e820), so the allocator's notion of
  "how much RAM exists" is honest on real hardware, not just qemu's
  128mb default.
- `user_shell.ld`/`user_prog.ld` collapse into one `user.ld` -- every
  process links at the same virtual base, since isolation no longer
  comes from address partitioning.
- `paging.nsc` (the scratch pool + `exec_shwin` remap trick) is
  retired entirely -- a nested `sh` just gets its own real address
  space like anything else, no special case.
 
## non-goals (explicitly deferred)
 
- **multicore/SMP** -- out of scope per the user, until needed.
- **SWORD** (a filesystem considered earlier) -- unrelated, deferred.
- **ring3/TSS/user-mode** -- a real, separate follow-up once this
  foundation exists.
- **a scheduler, TCBs, preemption, daemons** -- the very next project,
  built on top of this one. this spec does not add concurrency.
- **growing the identity map past 1GiB** -- usable RAM is capped at
  the first 1GiB (what `boot.s` already identity-maps) regardless of
  how much more e820 reports. a real machine with >1GiB works, it just
  doesn't get to use the excess yet -- a small, honest, separately-
  scoped follow-up, not silently bundled in here.
- **per-process stacks** -- execution stays synchronous/blocking (one
  thing running at a time, exactly like today), so the single shared
  kernel stack `boot.s` already set up remains correct and sufficient.
  real per-thread stacks become necessary once the scheduler project
  adds preemption, not before.
 
## architecture overview
 
new file pair: `src/kernel/vmm.nsc`/`vmm.nsh` (virtual memory manager),
matching the existing one-subsystem-per-file convention. it owns:
 
1. the physical frame bitmap allocator (`vmm_alloc_frame`/
   `vmm_free_frame`).
2. e820 memory-map ingestion (`vmm_init`, called from `kmain` in place
   of `paging_init`).
3. per-process page table construction/teardown (`vmm_create_addrspace`/
   `vmm_map_page`/`vmm_destroy_addrspace`).
 
`paging.nsc` is deleted. its only two real exports
(`paging_alloc_frame`/`paging_free_frame`) and the `exec_shwin_live`
mechanism in `exec.nsc` that calls them are deleted too -- superseded,
not kept alongside.
 
## the physical address layout (unchanged 1GiB identity map, re-split)
 
`boot.s`'s own 1GiB identity map is **not rebuilt** -- it's re-split,
by treating specific existing 2MiB `pd` slices differently once each
process gets its own copy of the `pd`:
 
| pd index | virtual range | today | under this design |
|---|---|---|---|
| 0-1 | `0x0`-`0x3fffff` | kernel image, boot's own tables, real-mode/BIOS area, the one shared kernel stack | **shared** -- copied by value into every process's own pd (same 2MiB huge-page entries, identical mapping) |
| 2-15 | `0x400000`-`0x1ffffff` (28MiB) | the two old windows (`sh`'s `0x400000`, everything else's `0x600000`-`0x2000000`) | **private** -- each process's own code/data/heap, backed by a per-process pt chain (4KiB pages), unmapped by default, filled in on demand |
| 16-23 | `0x2000000`-`0x2ffffff` (`heap_colosseum`, 16mib) | kernel heap | **shared** -- copied by value, unchanged |
| 24-511 | `0x3000000`-`0x3fffffff` | the old scratch-frame pool (`0x3000000`-`0x4000000`) + previously-unused space up to 1GiB | **general frame pool** -- what the new bitmap allocator draws 4KiB frames from, for both page-table-structure frames and process memory |
 
every process's own address space is therefore: a fresh pml4 (1 entry
populated, pointing at a fresh pdpt), a fresh pdpt (1 entry populated,
pointing at a fresh pd), and a fresh pd where indices 0-1 and 16-23 are
copied by value from the boot-time pd (still huge pages, still
identity) and indices 2-15 start absent, becoming present (pointing at
a freshly allocated per-process pt, 4KiB entries) only for whichever
2MiB slice(s) that process actually touches. a process that only needs
one 2MiB slice's worth of code+data+heap costs exactly one pt frame,
not fourteen.
 
*(the exact pd index boundaries above are read directly off today's
real addresses -- `heap_colosseum` at `0x2000000`/16mib,
`paging_pool_limit` at `0x4000000` -- confirm they haven't drifted
before implementing, same as any spec.)*
 
## physical frame allocator
 
a flat bitmap, one bit per 4KiB frame, sized for up to 4GiB of real
RAM (`2^32 / 4096 / 8` = 131072 bytes = 128KiB, a fixed static array in
`vmm.nsc` -- a real, generous, reasoned ceiling in the same spirit as
the 10-slot process table or klog's 8KiB buffer, not an arbitrary
guess). `vmm_alloc_frame()` linearly scans for a clear bit, sets it,
returns that frame's physical address (or `(u64)0`, the same
clean-failure sentinel every allocator in this kernel already uses, if
none are free). `vmm_free_frame(addr)` clears the bit. both are
`O(n)` over the bitmap -- matches the "ship the real minimum" standard
this kernel already holds itself to elsewhere; a free-list or other
`O(1)` scheme is a real, possible future improvement, not needed yet.
 
at `vmm_init()`:
- zero the whole bitmap (nothing free yet).
- for every e820 range marked "available," clear (mark free) the
  frames it covers, but never past the first 1GiB (see non-goals) and
  never below `0x3000000` (the general pool's own start -- everything
  below that is already spoken for per the layout table above and must
  never be handed out as a general-purpose frame).
- explicitly re-reserve (mark used) anything the e820 map might
  mistakenly call "available" but that's actually still in use below
  `0x3000000` -- belt-and-suspenders against a bad/incomplete map,
  matching the "trust but verify hardware-reported state" instinct
  `ata.nsc`'s own polling-timeout code already has.
 
## memory detection (e820)
 
two sources, converging on the same bitmap-init call:
 
- **qemu `-kernel` path (PVH):** `boot.s`'s own PVH entry already
  receives a `hvm_start_info` structure from the hypervisor, which
  includes a memory-map table (`memmap_paddr`/`memmap_entries`) per
  the PVH boot protocol -- `kmain` reads it directly, no BIOS call
  needed. *(confirm the exact field layout against `boot.s`'s current
  PVH entry code during implementation -- this spec describes the
  mechanism, not a verified-byte-for-byte struct offset.)*
- **real-BIOS path (`dynamite.s`):** a new real-mode `INT 15h,
  AX=E820h` loop, added to `dynamite.s` before the existing
  protected-mode transition, storing each returned entry into a fixed
  buffer `kmain` can read after the mode transition completes (the
  same "compute it once in real/32-bit mode, hand it forward" pattern
  `dynamite.s` already uses for `KERNEL_ENTRY_ADDR`/`KERNEL_LOAD_ADDR`,
  just data instead of assemble-time constants).
 
both paths write into one common, small, fixed-size array of
`(base, length, type)` entries `vmm_init()` reads identically
regardless of which boot path got the machine here -- same "both boot
paths converge on the identical `_start`" principle `kmain.nsc`'s own
top note already states for everything else.
 
## exec.nsc integration
 
`sys_exec`'s flow, with the two-window/`exec_shwin` logic removed and
vmm calls added:
 
1. resolve the path, read the file into `exec_filebuf` (unchanged).
2. `elf_validate` (unchanged). the old `elf_window`/`elf_window_lo`/
   `elf_window_hi` two-window lookup in `elf.nsc` is deleted --
   there's only one link convention now.
3. `vmm_create_addrspace()` -- allocates a fresh pml4/pdpt/pd (3
   frames), populates the pd's shared indices (0-1, 16-23) by copying
   the boot-time pd's own entries, leaves indices 2-15 absent. returns
   the new pml4's physical address, or `(u64)0` (the same
   clean-failure sentinel as `vmm_alloc_frame`/`arena_alloc`) if any of
   the 3 frames couldn't be allocated -- in which case whichever of the
   3 *did* succeed are freed again before returning, so a failed create
   never leaks frames.
4. walk the elf's `PT_LOAD` segments (same validation `elf_segments_ok`
   already does, minus the window-bounds check -- replaced by a
   simpler sanity check: `p_vaddr` falls inside the private range,
   `p_vaddr+p_memsz` doesn't overflow. the real ceiling is now "does
   the frame allocator have frames left," enforced naturally by
   `vmm_alloc_frame` returning 0, not a fixed window size). for each
   4KiB page a segment touches: `vmm_alloc_frame()`, `vmm_map_page()`
   into the new address space's private pd region (allocating a
   per-process pt on first use of a given 2MiB slice), then copy/zero
   the segment's bytes into that frame -- written directly through the
   *kernel's own* identity-mapped alias of that same physical address
   (no cr3 switch needed yet: every frame the allocator hands out
   comes from the general pool, which sits inside the 1GiB range the
   *kernel's own currently-active* page tables already identity-map).
5. `proc_push` (exec.nsc's existing 10-slot process table), extended
   with one more parallel array, `kaboom_proc_pml4[10]`, recording this
   process's own pml4 physical address alongside its pid/name.
6. switch `cr3` to the new address space.
7. `elf_call_entry` -- unchanged, still a blocking `call *rax`.
8. switch `cr3` back to the caller's own address space: for a nested
   exec (depth was already >= 1) that's `kaboom_proc_pml4` at the depth
   `proc_pop` is about to restore to; for the very first exec
   (`kmain`'s own boot-time `sys_exec("sh")`, depth 0 -> 1, no parent
   slot exists yet) that's the original boot-time `cr3` `boot.s` set
   up, saved once into a new global (`kernel_cr3` or similar) at
   `kmain`'s very start, before anything else runs.
9. `vmm_destroy_addrspace()` -- walks *only* the child's own private pd
   entries (2-15) and whatever per-process pt frames they point at,
   freeing every frame back to the bitmap, then frees the pd/pdpt/pml4
   frames themselves. the shared indices (0-1, 16-23) are never
   touched or freed -- they're the boot-time kernel's own memory, not
   this process's.
10. `proc_pop` (unchanged).
 
the boot-time `sh` (`kmain`'s own `sys_exec("sh", ...)` call) goes
through this *exact same* path -- no special case for "the first
process," matching this project's own stated goal of retiring the
`exec_shwin` special-casing rather than reintroducing a new one.
 
## link script unification
 
`user_shell.ld` and `user_prog.ld` collapse into one `user.ld`, linking
every program at `0x400000` (unchanged base, least disruptive) with no
upper bound tied to a shared window -- the real ceiling is now 28mib of
private virtual space (pd indices 2-15) per process, the same total
the single old shared window offered, just no longer contended.
 
## error handling
 
- **frame exhaustion during address-space construction** (partway
  through mapping `PT_LOAD` segments): `vmm_create_addrspace`'s whole
  construction is rolled back -- every frame already allocated for this
  half-built address space is freed, and `sys_exec` fails cleanly (the
  same ambiguous `-1` `sys_exec` already returns for "command not
  found" et al. -- consistent with `idt.nsc`'s own documented
  reasoning for why that's deliberately not more specific). matches
  `kfs`'s own "never write/allocate irreversibly until the whole
  operation is known to succeed" discipline.
- **frame exhaustion for an ordinary kernel-side `vmm_alloc_frame`
  call** (not process construction): returns `(u64)0`, the same
  clean-failure sentinel `arena_alloc` already established.
- **e820/PVH memory-map detection failing or returning nothing
  usable** at boot: a hard halt with `klog_write`, matching the
  `kfs_mount`/`arena_carve` halt pattern already established this
  session -- a boot-time invariant violation, not a runtime error a
  caller can recover from.
 
## testing plan
 
this replaces the *only* path anything executes through, so testing is
first and foremost a full regression pass, not just new-capability
testing:
 
- every existing coreutil (`cat`, `ls`, `cp`, `mv`, `chmod`, etc.)
  still works through the new per-process path.
- deep `sh` nesting (the 9-deep chain this session's proc-table fix
  proved) still works -- now via real, independent address spaces
  instead of the retired `exec_shwin` remap, and `/proc/ps` still
  reports correctly.
- shebang chains (the 8-deep limit) still work.
- **new**: two sequential execs, each writing a distinct value to the
  same virtual heap address, confirm the second never observes the
  first's data (proves real isolation, not an accidental leftover
  mapping).
- **new**: exhaust the frame pool deliberately (e.g. by shrinking it
  to a tiny test size) and confirm a subsequent exec fails cleanly with
  no corruption, then confirm a normal-sized exec succeeds again
  afterward (proves the rollback path actually returns every frame it
  took).
- all of the above via real qemu boot + serial interaction, per this
  session's established standard -- compiled and live-booted, not
  claimed from reading alone.
 
## open items / follow-ups (real, deliberately not attempted here)
 
- growing the identity map past 1GiB, so real machines with more RAM
  can use it.
- ring3/TSS -- real privilege separation, once this foundation exists.
- an `O(1)` frame allocator (free-list) if the linear bitmap scan ever
  actually shows up as a real cost.
- the scheduler + daemon project this one exists to unblock.
powered by btf.