| git.druid.rocks | index | druid520 | kaboom | docs/ | superpowers/ | specs/ | 2026-09-30-vmm-per-process-addrspace-design.md |
docs/superpowers/specs/2026-09-30-vmm-per-process-addrspace-design.md
# per-process address spaces (vmm)
## context / motivation
kaboom has no scheduler and no per-process memory isolation. every
"process" is an ordinary called function running in the single shared
address space `boot.s` built at boot: one identity-mapped 1GiB region
(`pml4[0] -> pdpt[0] -> pd[0..511]`, each pd entry a 2MiB huge page,
`pd[i]` covering physical/virtual `i*0x200000` .. `i*0x200000+0x1fffff`
-- identity, since phys==virt everywhere in this map). within that,
exactly two fixed windows exist for userland code: `0x400000` (`sh`
only, `user_shell.ld`) and `0x600000`-`0x2000000` (everything else,
`user_prog.ld`), each with **one** live occupant at a time --
`paging.nsc`'s scratch-frame pool + `exec_shwin` remap trick exists
solely to let a *nested* `sh` reuse the first window without colliding
with the outer one it's nested inside.
this is the direct blocker for real daemons: a daemon sitting in the
"everything else" window would prevent any other program from running
at all while it's alive. this project replaces the two-window scheme
with real, private, per-process address spaces, so an arbitrary number
of processes (bounded only by physical memory) can be resident at
once. it is a prerequisite for the scheduler + daemon project that
follows it -- that project adds concurrency (a timer, TCBs, preemption,
the daemon mechanism); this one only adds isolation. execution stays
exactly as synchronous/blocking as it is today (`sys_exec` still calls
`elf_call_entry` and doesn't return until the child does).
decisions already made with the user, carried into this spec:
- **ring0-only.** no TSS, no user/kernel privilege split. isolation
here means "processes can't collide with each other's memory by
accident," not "processes can't attack each other" -- ring3 is a
separate, later, well-scoped project (`gdt.nsc`'s own comment already
anticipates it).
- **real e820/memory-map detection**, not a hardcoded RAM size --
needed so the frame allocator never hands out a frame past the end
of real, installed memory (exactly the bug class `paging.nsc`'s own
history section already lived through once with the old scratch
pool).
- **4KiB pages for process-private memory** (not 2MiB huge pages) --
a 16KB program shouldn't burn a full 2MiB frame; on a 64-128MB
machine that's the difference between a handful of resident
processes and dozens.
## goals
- every `sys_exec`'d process gets its own private virtual region,
backed by physical frames nothing else can touch, allocated on
demand and fully reclaimed when the process exits.
- a general-purpose physical frame allocator (alloc *and* free, in any
order) replacing the old LIFO-only scratch pool.
- real physical memory detection (e820), so the allocator's notion of
"how much RAM exists" is honest on real hardware, not just qemu's
128mb default.
- `user_shell.ld`/`user_prog.ld` collapse into one `user.ld` -- every
process links at the same virtual base, since isolation no longer
comes from address partitioning.
- `paging.nsc` (the scratch pool + `exec_shwin` remap trick) is
retired entirely -- a nested `sh` just gets its own real address
space like anything else, no special case.
## non-goals (explicitly deferred)
- **multicore/SMP** -- out of scope per the user, until needed.
- **SWORD** (a filesystem considered earlier) -- unrelated, deferred.
- **ring3/TSS/user-mode** -- a real, separate follow-up once this
foundation exists.
- **a scheduler, TCBs, preemption, daemons** -- the very next project,
built on top of this one. this spec does not add concurrency.
- **growing the identity map past 1GiB** -- usable RAM is capped at
the first 1GiB (what `boot.s` already identity-maps) regardless of
how much more e820 reports. a real machine with >1GiB works, it just
doesn't get to use the excess yet -- a small, honest, separately-
scoped follow-up, not silently bundled in here.
- **per-process stacks** -- execution stays synchronous/blocking (one
thing running at a time, exactly like today), so the single shared
kernel stack `boot.s` already set up remains correct and sufficient.
real per-thread stacks become necessary once the scheduler project
adds preemption, not before.
## architecture overview
new file pair: `src/kernel/vmm.nsc`/`vmm.nsh` (virtual memory manager),
matching the existing one-subsystem-per-file convention. it owns:
1. the physical frame bitmap allocator (`vmm_alloc_frame`/
`vmm_free_frame`).
2. e820 memory-map ingestion (`vmm_init`, called from `kmain` in place
of `paging_init`).
3. per-process page table construction/teardown (`vmm_create_addrspace`/
`vmm_map_page`/`vmm_destroy_addrspace`).
`paging.nsc` is deleted. its only two real exports
(`paging_alloc_frame`/`paging_free_frame`) and the `exec_shwin_live`
mechanism in `exec.nsc` that calls them are deleted too -- superseded,
not kept alongside.
## the physical address layout (unchanged 1GiB identity map, re-split)
`boot.s`'s own 1GiB identity map is **not rebuilt** -- it's re-split,
by treating specific existing 2MiB `pd` slices differently once each
process gets its own copy of the `pd`:
| pd index | virtual range | today | under this design |
|---|---|---|---|
| 0-1 | `0x0`-`0x3fffff` | kernel image, boot's own tables, real-mode/BIOS area, the one shared kernel stack | **shared** -- copied by value into every process's own pd (same 2MiB huge-page entries, identical mapping) |
| 2-15 | `0x400000`-`0x1ffffff` (28MiB) | the two old windows (`sh`'s `0x400000`, everything else's `0x600000`-`0x2000000`) | **private** -- each process's own code/data/heap, backed by a per-process pt chain (4KiB pages), unmapped by default, filled in on demand |
| 16-23 | `0x2000000`-`0x2ffffff` (`heap_colosseum`, 16mib) | kernel heap | **shared** -- copied by value, unchanged |
| 24-511 | `0x3000000`-`0x3fffffff` | the old scratch-frame pool (`0x3000000`-`0x4000000`) + previously-unused space up to 1GiB | **general frame pool** -- what the new bitmap allocator draws 4KiB frames from, for both page-table-structure frames and process memory |
every process's own address space is therefore: a fresh pml4 (1 entry
populated, pointing at a fresh pdpt), a fresh pdpt (1 entry populated,
pointing at a fresh pd), and a fresh pd where indices 0-1 and 16-23 are
copied by value from the boot-time pd (still huge pages, still
identity) and indices 2-15 start absent, becoming present (pointing at
a freshly allocated per-process pt, 4KiB entries) only for whichever
2MiB slice(s) that process actually touches. a process that only needs
one 2MiB slice's worth of code+data+heap costs exactly one pt frame,
not fourteen.
*(the exact pd index boundaries above are read directly off today's
real addresses -- `heap_colosseum` at `0x2000000`/16mib,
`paging_pool_limit` at `0x4000000` -- confirm they haven't drifted
before implementing, same as any spec.)*
## physical frame allocator
a flat bitmap, one bit per 4KiB frame, sized for up to 4GiB of real
RAM (`2^32 / 4096 / 8` = 131072 bytes = 128KiB, a fixed static array in
`vmm.nsc` -- a real, generous, reasoned ceiling in the same spirit as
the 10-slot process table or klog's 8KiB buffer, not an arbitrary
guess). `vmm_alloc_frame()` linearly scans for a clear bit, sets it,
returns that frame's physical address (or `(u64)0`, the same
clean-failure sentinel every allocator in this kernel already uses, if
none are free). `vmm_free_frame(addr)` clears the bit. both are
`O(n)` over the bitmap -- matches the "ship the real minimum" standard
this kernel already holds itself to elsewhere; a free-list or other
`O(1)` scheme is a real, possible future improvement, not needed yet.
at `vmm_init()`:
- zero the whole bitmap (nothing free yet).
- for every e820 range marked "available," clear (mark free) the
frames it covers, but never past the first 1GiB (see non-goals) and
never below `0x3000000` (the general pool's own start -- everything
below that is already spoken for per the layout table above and must
never be handed out as a general-purpose frame).
- explicitly re-reserve (mark used) anything the e820 map might
mistakenly call "available" but that's actually still in use below
`0x3000000` -- belt-and-suspenders against a bad/incomplete map,
matching the "trust but verify hardware-reported state" instinct
`ata.nsc`'s own polling-timeout code already has.
## memory detection (e820)
two sources, converging on the same bitmap-init call:
- **qemu `-kernel` path (PVH):** `boot.s`'s own PVH entry already
receives a `hvm_start_info` structure from the hypervisor, which
includes a memory-map table (`memmap_paddr`/`memmap_entries`) per
the PVH boot protocol -- `kmain` reads it directly, no BIOS call
needed. *(confirm the exact field layout against `boot.s`'s current
PVH entry code during implementation -- this spec describes the
mechanism, not a verified-byte-for-byte struct offset.)*
- **real-BIOS path (`dynamite.s`):** a new real-mode `INT 15h,
AX=E820h` loop, added to `dynamite.s` before the existing
protected-mode transition, storing each returned entry into a fixed
buffer `kmain` can read after the mode transition completes (the
same "compute it once in real/32-bit mode, hand it forward" pattern
`dynamite.s` already uses for `KERNEL_ENTRY_ADDR`/`KERNEL_LOAD_ADDR`,
just data instead of assemble-time constants).
both paths write into one common, small, fixed-size array of
`(base, length, type)` entries `vmm_init()` reads identically
regardless of which boot path got the machine here -- same "both boot
paths converge on the identical `_start`" principle `kmain.nsc`'s own
top note already states for everything else.
## exec.nsc integration
`sys_exec`'s flow, with the two-window/`exec_shwin` logic removed and
vmm calls added:
1. resolve the path, read the file into `exec_filebuf` (unchanged).
2. `elf_validate` (unchanged). the old `elf_window`/`elf_window_lo`/
`elf_window_hi` two-window lookup in `elf.nsc` is deleted --
there's only one link convention now.
3. `vmm_create_addrspace()` -- allocates a fresh pml4/pdpt/pd (3
frames), populates the pd's shared indices (0-1, 16-23) by copying
the boot-time pd's own entries, leaves indices 2-15 absent. returns
the new pml4's physical address, or `(u64)0` (the same
clean-failure sentinel as `vmm_alloc_frame`/`arena_alloc`) if any of
the 3 frames couldn't be allocated -- in which case whichever of the
3 *did* succeed are freed again before returning, so a failed create
never leaks frames.
4. walk the elf's `PT_LOAD` segments (same validation `elf_segments_ok`
already does, minus the window-bounds check -- replaced by a
simpler sanity check: `p_vaddr` falls inside the private range,
`p_vaddr+p_memsz` doesn't overflow. the real ceiling is now "does
the frame allocator have frames left," enforced naturally by
`vmm_alloc_frame` returning 0, not a fixed window size). for each
4KiB page a segment touches: `vmm_alloc_frame()`, `vmm_map_page()`
into the new address space's private pd region (allocating a
per-process pt on first use of a given 2MiB slice), then copy/zero
the segment's bytes into that frame -- written directly through the
*kernel's own* identity-mapped alias of that same physical address
(no cr3 switch needed yet: every frame the allocator hands out
comes from the general pool, which sits inside the 1GiB range the
*kernel's own currently-active* page tables already identity-map).
5. `proc_push` (exec.nsc's existing 10-slot process table), extended
with one more parallel array, `kaboom_proc_pml4[10]`, recording this
process's own pml4 physical address alongside its pid/name.
6. switch `cr3` to the new address space.
7. `elf_call_entry` -- unchanged, still a blocking `call *rax`.
8. switch `cr3` back to the caller's own address space: for a nested
exec (depth was already >= 1) that's `kaboom_proc_pml4` at the depth
`proc_pop` is about to restore to; for the very first exec
(`kmain`'s own boot-time `sys_exec("sh")`, depth 0 -> 1, no parent
slot exists yet) that's the original boot-time `cr3` `boot.s` set
up, saved once into a new global (`kernel_cr3` or similar) at
`kmain`'s very start, before anything else runs.
9. `vmm_destroy_addrspace()` -- walks *only* the child's own private pd
entries (2-15) and whatever per-process pt frames they point at,
freeing every frame back to the bitmap, then frees the pd/pdpt/pml4
frames themselves. the shared indices (0-1, 16-23) are never
touched or freed -- they're the boot-time kernel's own memory, not
this process's.
10. `proc_pop` (unchanged).
the boot-time `sh` (`kmain`'s own `sys_exec("sh", ...)` call) goes
through this *exact same* path -- no special case for "the first
process," matching this project's own stated goal of retiring the
`exec_shwin` special-casing rather than reintroducing a new one.
## link script unification
`user_shell.ld` and `user_prog.ld` collapse into one `user.ld`, linking
every program at `0x400000` (unchanged base, least disruptive) with no
upper bound tied to a shared window -- the real ceiling is now 28mib of
private virtual space (pd indices 2-15) per process, the same total
the single old shared window offered, just no longer contended.
## error handling
- **frame exhaustion during address-space construction** (partway
through mapping `PT_LOAD` segments): `vmm_create_addrspace`'s whole
construction is rolled back -- every frame already allocated for this
half-built address space is freed, and `sys_exec` fails cleanly (the
same ambiguous `-1` `sys_exec` already returns for "command not
found" et al. -- consistent with `idt.nsc`'s own documented
reasoning for why that's deliberately not more specific). matches
`kfs`'s own "never write/allocate irreversibly until the whole
operation is known to succeed" discipline.
- **frame exhaustion for an ordinary kernel-side `vmm_alloc_frame`
call** (not process construction): returns `(u64)0`, the same
clean-failure sentinel `arena_alloc` already established.
- **e820/PVH memory-map detection failing or returning nothing
usable** at boot: a hard halt with `klog_write`, matching the
`kfs_mount`/`arena_carve` halt pattern already established this
session -- a boot-time invariant violation, not a runtime error a
caller can recover from.
## testing plan
this replaces the *only* path anything executes through, so testing is
first and foremost a full regression pass, not just new-capability
testing:
- every existing coreutil (`cat`, `ls`, `cp`, `mv`, `chmod`, etc.)
still works through the new per-process path.
- deep `sh` nesting (the 9-deep chain this session's proc-table fix
proved) still works -- now via real, independent address spaces
instead of the retired `exec_shwin` remap, and `/proc/ps` still
reports correctly.
- shebang chains (the 8-deep limit) still work.
- **new**: two sequential execs, each writing a distinct value to the
same virtual heap address, confirm the second never observes the
first's data (proves real isolation, not an accidental leftover
mapping).
- **new**: exhaust the frame pool deliberately (e.g. by shrinking it
to a tiny test size) and confirm a subsequent exec fails cleanly with
no corruption, then confirm a normal-sized exec succeeds again
afterward (proves the rollback path actually returns every frame it
took).
- all of the above via real qemu boot + serial interaction, per this
session's established standard -- compiled and live-booted, not
claimed from reading alone.
## open items / follow-ups (real, deliberately not attempted here)
- growing the identity map past 1GiB, so real machines with more RAM
can use it.
- ring3/TSS -- real privilege separation, once this foundation exists.
- an `O(1)` frame allocator (free-list) if the linear bitmap scan ever
actually shows up as a real cost.
- the scheduler + daemon project this one exists to unblock.