do not edit — generated by btf.
git.druid.rocksindexdruid520kaboomdocs/kernel.btft

docs/kernel.btft


[link rel="stylesheet" href="keyframes.css"][e]
 
[table class="topnav"]
[tr]
[td class="logotab"][see name="index"]kaboom[e][e]
[td][see name="kernel"]kernel[e][e]
[td][see name="fs"]fs[e][e]
[td][see name="syscalls"]syscalls[e][e]
[td][see name="userland"]userland[e][e]
[td][see name="build"]build[e][e]
[e]
[e]
 
[h1]kernel[e]
 
this page is the boot-to-running-process story: how a multiboot loader's 32-bit protected-mode handoff turns into a 64-bit kernel that can load and run a real elf binary, plus the drivers that make the machine feel like a machine while that's happening. read [see name="syscalls"]syscalls[e] for the actual syscall table -- this page only covers the dispatch mechanism, not the 24 individual calls.
 
[h2]boot: identity paging before anything else[e]
 
[tt]src/boot/boot.s[e] is 32-bit hand-written asm, not nsc, because nothing nsc-compiled can run yet: nscc only emits 64-bit object files, and long mode itself has a hard requirement that paging be turned on before you're allowed into it. get any part of this wrong -- a bad page table entry, cr4/cr3/efer set in the wrong order -- and the cpu triple-faults silently. no error, no message, just an instant reboot. that makes this the single highest-risk file in the whole kernel, and it's tested in isolation before anything else gets built on top of it.
 
the loader story itself is a real qemu-specific wart, worth knowing about if you ever try to boot this on anything else: qemu's [tt]-kernel[e] multiboot1 path hard-refuses any ELFCLASS64 image outright ("Cannot load x86-64 image, give a 32bit one"), and that check runs before qemu even looks at what's inside the load segments. a multiboot1 header alongside a PVH note was tried, and it doesn't work -- qemu detects multiboot1 first and its 32-bit check fires unconditionally, so you get one or the other, never both. PVH (the xen/hvm direct-boot protocol, which qemu also implements) is fine with an ELF64 container; it just needs a note pointing at a 32-bit physical entry address. that's the [tt].note.pvh[e] section at the top of the file, and it's still the fast dev-loop path today, completely unchanged.
 
a genuinely real-hardware-bootable path exists now too, though it isn't a multiboot1 header -- it never could be, given the constraint above -- it's a second, independent bootloader ([tt]src/boot/dynamite.s[e], see [see name="build"]build[e]'s own coverage of [tt]mk/bl.sh[e]) that a real BIOS actually understands: raw sector 0, real mode, no elf notes involved at all. both boot paths converge on the identical [tt]_start[e] in the identical cpu state PVH already guaranteed -- almost unchanged: this file gained exactly one addition for the new path, a [tt].bss[e]-zeroing loop right at the top of [tt]_start[e], since real hardware ram doesn't start zeroed the way qemu's guest ram happens to, and this file's own page tables (and any zero-initialized kernel global) depend on actually being zero. an elf loader already does that for free, which is why the PVH path never needed it. everything else below in this section runs identically regardless of which path got the machine here.
 
[tt]_start[e] builds one pml4 entry -> one pdpt entry -> 512 2mib huge pages in [tt]pd[e], identity-mapping the first 1gib (virtual == physical everywhere). that's a lot more than the kernel needs at boot, but it means this table doesn't need revisiting for a good while -- and it turns out that decision pays off directly: [tt]vmm.nsc[e] (below) copies this table's own entries straight into every fresh per-process address space by value, never needs to build the shared mapping from scratch itself. after the fill loop: CR4.PAE, CR3 loaded with the pml4 base, EFER.LME via the MSR, then CR0.PG -- and only after all four of those does the code far-jump into the 64-bit code segment ([tt]long_mode_entry[e]). the very first thing that segment does, before calling [tt]kmain[e], is write a raw [tt]'B'[e] directly to COM1 (port 0x3f8) with a bare [tt]outb[e], independent of any nsc code or the [tt]io.s[e] helpers being correct -- if that byte never shows up in the serial log, the fault is in this file, not in anything kmain calls. it's a sanity beacon, not a real driver, and it exists specifically because a triple fault here gives you nothing else to go on.
 
[h2]gdt / idt: segments, interrupts, syscalls[e]
 
[tt]gdt.nsc[e] replaces boot.s's own hand-built [tt]gdt64[e] (which only exists to get the cpu into long mode at all) with one the kernel actually owns: null, kernel code, kernel data. the reason to bother once boot.s already has a working gdt: this one has somewhere real to grow -- a tss entry once [tt]idt.nsc[e] wants a dedicated interrupt stack, or user-mode segments once elf programs actually run in ring 3, neither of which exists yet. gdt entries can't be an nsc [tt]struct[e], because nsc struct fields are word-addressed rather than byte-packed, so each 8-byte entry gets built by hand (base/limit/access/granularity packed into a [tt]u64[e]) and written with a single 8-byte [tt]*ptr[e] store. [tt]idt.nsc[e] hits the exact same packing constraint for its 16-byte entries.
 
the idt sets up 32 cpu exception vectors, remaps the pic's 16 irq vectors up to 32-47 (with everything except irq1/keyboard and irq4/com1 masked off -- nothing else has a driver hooked up yet), and one syscall gate at vector 0x80 with dpl=3 so user mode will eventually be able to [tt]int $0x80[e] from ring 3. every vector funnels through one shared C-side dispatcher, [tt]isr_dispatch[e], called by a single shared asm trampoline in [tt]idt_asm.s[e]. vector 128 goes straight to [tt]syscall_dispatch[e]; vectors under 32 are real cpu exceptions and kaboom has no recovery story for those yet -- it prints the vector/error code to serial and halts forever; anything else is a remapped hardware irq, which always gets an eoi sent back to the pic (or it stops delivering further interrupts entirely -- an easy thing to get wrong, since nothing breaks until the [i]second[e] keypress, not the first) and then gets handed to whichever driver owns it (irq1 -> keyboard, irq4 -> com1/serial).
 
the syscall convention itself: number in rax (also where the return value goes), args in rdi/rsi/rdx, kept in the same register order/positions as the real x86-64 sysv/syscall convention purely for familiarity -- even though the actual entry mechanism is the classic int-gate here, not the [tt]syscall[e] instruction. every syscall that can genuinely fail routes its result through one shared function, [tt]sys_doerror[e], instead of hand-rolling "log it, then set rax" at each call site. [tt]sys_doerror[e] takes the caller's own already-correct "did this fail" test (the failure convention genuinely differs per syscall -- some return negative, some a plain 0/1 -- so it never tries to guess that from the result value alone), and if it did fail, writes a kernel log line: [tt]"pid <N>: <op>: permission denied"[e] when [tt]kaboom_errno[e] says that's really why, or a plainer [tt]"<op>: failed"[e] otherwise. [tt]kaboom_errno[e] (set in [tt]kfs.nsc[e]) has exactly one place that ever sets it to 1 -- [tt]kfs_check_perm[e] -- and every other public entry point with its own permission check resets it to 0 at its own start, so a stale value from an earlier, unrelated failure can never leak into a later one's log line. [tt]klogs[e] shows these lines like any other kernel log entry. the full syscall table -- all 26 calls, their exact signatures, what each hands back on failure -- lives on [see name="syscalls"]syscalls[e], generated straight from [tt]idt.nsc[e]'s own numbered doc comment so the two can't quietly drift apart; this page only covers the mechanism.
 
[h2]alloc: the kernel heap[e]
 
[tt]alloc.nsc[e] is a bump allocator, nothing more -- one parent arena ([tt]heap_arena[e]) with named child arenas ([tt]fs_arena[e], 4mib) carved out of its remaining space, [tt]arena_alloc[e] just advances a pointer and never frees, the same tradeoff every allocator in this kernel makes (kalloc/[tt]sys_alloc[e], [tt]fs_alloc[e], the userspace [tt]malloc[e] built on top -- see [see name="userland"]userland[e]). [tt]kalloc_used[e]/[tt]kalloc_total[e]/[tt]fs_alloc_used[e]/[tt]fs_alloc_total[e] are the live numbers [tt]/int/mem[e] ([see name="fs"]fs[e]'s virtfs coverage) reports back.
 
[tt]heap_arena[e] starts at a fixed [tt]0x2000000[e] (32mib) now, 16mib wide -- it didn't always, and that was a real, severe, live-reproduced bug: it used to start right at [tt]get_kernel_end()[e] (roughly [tt]0x142000[e]), which put its own 16mib span squarely on top of [i]both[e] fixed userspace program-load addresses (sh at [tt]0x400000[e], everything else at [tt]0x600000[e] -- see [tt]user_shell.ld[e]/[tt]user_prog.ld[e]). since [tt]arena_alloc[e] never frees, past roughly 2.7mib of cumulative allocation in one boot, new heap allocations silently started landing on running program memory instead. this was invisible for a long time because it took a lot of small allocations to reach that point -- until whole-file reads sized to the real file (see [see name="userland"]userland[e]'s [tt]read_all[e] coverage) made a single large allocation routine instead of rare: roughly a dozen large file reads in one boot was enough to reach the same 2.7mib line that used to take dozens of small ones. reproduced directly: repeated large reads eventually panicked the kernel with sh's own code overwritten (an invalid-opcode fault), and once, worse, corrupted the on-disk kfs superblock itself through a stray write via a pointer that had wandered into program memory -- the filesystem wouldn't mount on the next boot at all. the fix is the fixed [tt]0x2000000[e] start above: comfortably clear of both program windows, comfortably inside boot.s's own 1gib identity map, comfortably within [tt]mk.conf[e]'s [tt]-m 128[e] once [tt]fs_arena[e]'s own carve-out and [tt]vmm.nsc[e]'s own general frame allocator (which starts handing out frames exactly at [tt]heap_arena[e]'s own end, see below) are both accounted for. [tt]fs_arena[e] was widened from 1mib to 4mib in the same change, now that the heap has real room to spare -- it caps how big a single open file can be (see [see name="fs"]fs[e]'s honest accounting of what's actually reachable versus the format's own ceiling).
 
[h2]vmm: physical frames and real per-process address spaces[e]
 
[tt]vmm.nsc[e] is a genuine virtual memory subsystem now, not the narrow one-trick patch that used to sit here: every process gets its own real address space, sh included, with no special-casing for sh's old self-nesting problem (a shebang script naming [tt]sh[e] as its own interpreter, exec'd from within an already-running sh). that problem doesn't need a trick anymore -- a nested sh's whole address space is already a different physical page table from the outer one's, no different from how any other process's exec works.
 
[tt]vmm_init[e] turns the real memory map handed up from the boot path (pvh's own [tt]hvm_start_info[e], or dynamite's real int 15h/e820h call -- see [tt]vmm_e820_from_pvh[e]/[tt]vmm_e820_from_bios[e]) into a flat, 4kib-granular bitmap ([tt]vmm_frame_bitmap[e]), capped at the first 1gib the same way everything else here still is. everything at or below [tt]heap_arena[e]'s own end (the kernel image, boot's own page tables, the stack, [tt]heap_arena[e] itself) starts marked used and is never handed out as a general frame; [tt]vmm_alloc_frame[e]/[tt]vmm_free_frame[e] scan and flip single bits in what's left -- a real allocate-and-free pair, not a bump-only pool and not a lifo-only release scheme, since a process's whole address space, however many frames it ends up owning, gets freed in one shot when it exits rather than one frame at a time in careful reverse order.
 
[tt]vmm_create_addrspace[e] builds a fresh, empty pml4/pdpt/pd (three freshly allocated frames, zeroed) for a new process, then copies the kernel's own shared mappings into it by value straight off the boot-time page directory ([tt]get_pd_table[e], [tt]vmm_asm.s[e]): low memory/stack, plus [tt]heap_arena[e] and the entire general frame pool above it. the one range deliberately left unmapped is the private region every process gets to itself, [tt]0x400000[e]-[tt]0x2000000[e] -- [tt]vmm_map_page[e] fills that in lazily, one 4kib page at a time, as [tt]exec.nsc[e] (below) maps and copies each [tt]PT_LOAD[e] segment of whatever's being loaded. [tt]vmm_destroy_addrspace[e] walks that same private region back down once the process is done, frees every frame it finds mapped there plus the pml4/pdpt/pd themselves, and leaves the shared kernel entries alone -- they were never this process's to free.
 
this replaced a real, found-live bug on the way in: the first command run from a freshly-booted shell page-faulted while its own address space was still being built, because the shared-mapping copy above only covered the heap's own narrow footprint at first, not the whole range the general frame pool can actually hand frames out of -- any pml4/pdpt/pd/pt scratch frame (or a [tt]PT_LOAD[e] segment's own backing page) allocated from further out in the pool got dereferenced as a bare physical==virtual pointer by kernel code running under the [i]new[e], not-yet-fully-built address space, and faulted the instant it landed outside that narrower copied range. widening the copy to cover the pool's entire possible span -- cheap, a few hundred more quad writes, nothing boot.s's own page directory doesn't already have present there regardless of how much real ram is actually installed -- fixed it for good.
 
worth being explicit about what this still deliberately does NOT do: no page permissions beyond present+writable (every mapping here, shared or private, carries the same flags -- no nx, no read-only, exactly like boot.s's own identity map always has), no demand paging or swapping beyond the one-time lazy fill [tt]exec.nsc[e] already does while loading, no memory past the first 1gib (the same e820-aware-but-still-capped posture the fixed-ceiling pool before it already had). what's genuinely new: separate physical frames and separate page tables per process, torn down wholesale on exit instead of trickled back one frame at a time -- "extremely simple" isn't the right description for this file anymore, but it's still the minimum that actually fixes the problem it was built for, not a general-purpose mm.
 
and worth being equally honest about what "separate physical frames and separate page tables per process" does and doesn't actually guarantee: this is not hardware-enforced isolation, and was never meant to be. kaboom has no ring3, no tss, no permission enforcement of any kind -- everything runs in ring0 throughout (a confirmed, deliberate decision, not a gap waiting to be closed), and every process's own address space identity-maps the ENTIRE shared pd range by value (indices 16-511, physical -- most of the first 1gib, see [tt]vmm_create_addrspace[e] above). nothing stops ring0 code -- which is all of it -- from directly dereferencing another live process's private frames, or hand-building its own page table entries to reach anywhere it likes; there's no permission bit and no privilege level standing in the way. what this design actually delivers is "processes don't collide with each other BY ACCIDENT" (the real, stated goal this whole task existed for), not "processes can't deliberately read or corrupt each other" (never a goal here, and not implemented). relatedly, the userspace heap ([tt]sys_alloc[e], backed by the one shared [tt]heap_colosseum[e]) is NOT part of a process's private region at all -- it's ordinary shared memory, reachable from every address space exactly like the rest of indices 16-511, and like everything else [tt]heap_colosseum[e] ever hands out, it's never reclaimed when the process that allocated it exits. that's a real gap against the design spec's own "code/data/heap" framing of what's private per-process: only the code+data a [tt]PT_LOAD[e] segment actually maps is private and torn down on exit; heap allocations outlive their process forever, the same leak this kernel has always had. one concrete, if modest, cost of the private-region design worth being upfront about: physical [tt]0x400000[e]-[tt]0x2000000[e] (28mib) is now permanently excluded from the general frame pool ([tt]vmm_alloc_frame[e] never hands it out) -- it can't be identity-mapped the same way in every address space (that's exactly the range each process's own private region needs to differ in) while also being a shared, generally-allocatable frame, so it's simply carved out of the pool entirely, the same way [tt]heap_colosseum[e]'s own span always was.
 
[h2]exec: loading and running a program[e]
 
[tt]exec.nsc[e] is syscall 7, and it's the thing a shell actually needs: it wraps [tt]kfs_dir_find[e]/[tt]kfs_read_file[e]/[tt]elf_load[e]/[tt]elf_call_entry[e] -- four kernel-internal functions kmain's own test code had already used -- behind one call userspace can make. [tt]elf_call_entry[e]'s own asm does [tt]call *rax; ret[e], so whatever the loaded program leaves in [tt]%rax[e] when it returns becomes [tt]sys_exec[e]'s own return value too: a c-style [tt]int main(void)[e] return value doubling as an exit code, since there's no separate exit syscall (see [tt]src/user/crt.s[e]'s note on why returning already is exiting).
 
path resolution goes through [tt]kfs_resolve[e], not a bare-name-only lookup, so a real absolute or relative multi-component path ([tt]/bin/sh[e], [tt]a/b[e]) resolves correctly -- needed for conventional shebang lines like [tt]#!/bin/sh[e] to work at all. if that fails, there's a [tt]/bin[e] [tt]$PATH[e] fallback: [tt]kfs_dir_find_in[e] against [tt]kfs_bin_dir_lba[e] (resolved once at mount time), the same way a bare command name typed at the shell finds its binary. that fallback used to be a real, if mostly harmless, bug: the bare name that resolved via the fallback wasn't independently resolvable by anything downstream that didn't also know about [tt]kfs_bin_dir_lba[e] -- so a shebang script only reachable through the [tt]/bin[e] fallback would get passed to its interpreter as an unresolvable bare name, and the interpreter's own plain [tt]sys_open[e] on it would fail with a misleading "script not found", indistinguishable from the script genuinely not existing. the fix builds a real, independently-resolvable [tt]/bin/<name>[e] path with [tt]kfs_join_path[e] and uses that everywhere downstream instead -- for the interpreter's own [tt]argv[e] slot 1 and for the process-table entry both. [tt]kfs_join_path[e] itself replaced code that used to hand-spell the five ascii bytes of [tt]"/bin/"[e] as individual numbered constants, the one place in this kernel that hardcoded a path as raw bytes instead of a string literal -- not just uglier, but a second, more fragile source of truth for the same "bin" name [tt]kfs_mount[e] already had in a normal string. [tt]kfs_join_path[e] is now a genuinely reusable primitive, not something specific to this one call site.
 
permission enforcement happens right after resolution: [tt]kfs_check_perm(n, 4)[e] (the execute bit) has to pass before anything gets read or run, and that check applies identically to a shebang interpreter's own re-exec as it does to the original script or binary -- there's no separate, weaker path for "the thing sys_exec found on its own."
 
the shebang mechanism itself: if the resolved file isn't a valid elf but starts with [tt]#![e], [tt]shebang_parse_interp[e] reads the interpreter path off the first line (no shebang-line arguments, e.g. no [tt]#!/bin/sh -x[e] -- "super simple, just commands"), and [tt]sys_exec[e] re-execs that interpreter recursively, with argv shifted the same way a real unix [tt]execve()[e] does: interpreter, then the script's own path, then whatever args the original call had past its own [tt]argv[e] slot 0. kaboom's only real interpreter is sh itself, and a shebang naming sh gets exec'd from [i]within[e] an already-running sh -- that recursive call is no special case anymore: it gets a genuinely fresh address space exactly like any other exec does (see [tt]vmm.nsc[e], above), built by [tt]vmm_create_addrspace[e] and mapped page by page via [tt]vmm_map_page[e] before the nested interpreter ever runs, then torn down by [tt]vmm_destroy_addrspace[e] once that recursive call returns and the outer sh's own mapping is live again via [tt]load_cr3[e]. a hard nesting limit ([tt]exec_shebang_depth[e], 8) still fails the exec cleanly rather than ever letting a chain nest deeper -- not because a real per-process frame pool could run dry (it's thousands of frames deep now, nowhere near exhausted by 8 levels), but because nothing this shallow ever needs to nest further, and the limit closes the same self-referential-shebang stack overflow it always did (see below).
 
there used to be no such limit at all, on the theory that a shebang chain pointing back to itself would just "recurse until the shared 64kib stack runs out" -- which is exactly what it did, for real: a script whose own shebang line named itself ran the shared kernel stack ([tt]boot.s[e]) straight down into the live page tables sitting directly below it in memory, with no guard page of any kind in between, and triple-faulted the machine. the depth-8 limit closes the actual trigger directly (8 is far deeper than any real script chain nests, and matches the frame pool's own size exactly, so the pool can never run dry from legitimate nesting either); a second, independent line of defense went into [tt]boot.s[e] itself -- a generous, several-hundred-kib [tt].skip[e] padding gap between the page tables and the stack, pure [tt].bss[e] so it costs nothing real, there specifically so [i]any[e] deep stack use (not just this one, now-bounded, trigger) has to overflow much further before reaching anything live. that's defense in depth, not a structural guarantee -- a genuine unmapped guard page below the stack (a real, not-present 4kib page, so an overflow faults immediately instead of corrupting anything) would need splitting [tt]boot.s[e]'s own first 2mib huge page into a real 4kib page table, a bigger change than either of these two fixes and a real future candidate, not attempted here.
 
the process table this all threads through ([tt]kaboom_proc_depth[e] plus a real 10-slot [tt]kaboom_proc_pid[e]/[tt]kaboom_proc_name[e] array) is the honest answer to "what even is a process here": kaboom has no pcb and no scheduler. the process [i]is[e] the call stack -- [tt]sys_exec[e] doesn't return to its own caller until the child returns, and [tt]elf_call_entry[e] does a literal [tt]call *rax[e]. that's still a real, if shallow, parent/child tree: kmain's own boot-time [tt]sys_exec("sh")[e] is the first process, and every command sh's shell loop execs is that process's child. only sh ever calls exec again -- but only sh doing it doesn't bound depth at 2: a nested sh (the user typing [tt]sh[e], or a [tt]sh script[e] shebang chain) can itself exec a further command, and a table sized for depth 2 used to silently lose that nested sh's own identity the moment it did (found live: [tt]sh[e] then [tt]cat /proc/ps[e] from inside it showed only the outer sh and [tt]cat[e] -- the nested sh's own row never appeared, overwritten in the one "child" slot the table had past the first). the real ceiling isn't tied to any physical resource anymore, though: the process table itself (10 slots: [tt]kaboom_proc_pid[e]/[tt]kaboom_proc_name[e]/[tt]kaboom_proc_pml4[e]) is the actual limit, and [tt]sys_exec[e] checks [tt]kaboom_proc_depth[e] against it explicitly before ever pushing, since [tt]proc_push[e] itself has no bounds check of its own and a 10-deep stack really is this table's whole real capacity. that didn't used to be the thing that mattered: under the old shared-window design, a much smaller physical frame pool (8 frames) always ran out and failed the exec first, long before the table itself could overflow; the real per-process frame pool [tt]vmm.nsc[e] uses now is easily deep enough to nest straight past 10 without ever running dry, so this explicit depth check is what has to catch it instead -- confirmed for real: without it, nesting [tt]sh[e] past depth 10 corrupted memory silently (no panic, the machine just stopped responding a couple of levels further in) rather than failing the exec cleanly. [tt]/proc/self[e] and [tt]/proc/ps[e] ([tt]virtfs.nsc[e]) read this table directly, [tt]/proc/ps[e] listing every populated slot rather than assuming there are at most two.
 
there's no fd inheritance or passing here either, and [tt]sys_exec[e] now enforces that instead of just assuming it: it snapshots which of the 5 file-descriptor slots ([tt]fd_open_mask[e], [see name="fs"]fs[e]/[tt]fd.nsc[e]) are open right before calling the loaded program's entry point, and closes ([tt]fd_close_unless[e]) anything open at return that wasn't open before -- since nothing a child opens can ever legitimately outlive it, whatever's newly open at that point is the child's own leak, not something the caller still needs. a program that opened a file and forgot to close it used to leak that slot permanently, and only 5 exist -- a handful of buggy execs in one boot session used to be able to exhaust all of them until reboot. sh itself, the one process that stays resident across nested execs, keeps everything it already had open, since the snapshot is taken fresh at each new exec, not once at boot.
 
[h2]elf: what actually gets parsed[e]
 
[tt]elf.nsc[e] is a minimal elf64/x86-64 loader: [tt]elf_validate[e] checks the magic bytes, the ELFCLASS64 class byte, and the [tt]EM_X86_64[e] machine field, and [tt]elf_segments_ok[e] checks every [tt]PT_LOAD[e] segment's [tt]p_vaddr[e]/[tt]p_filesz[e]/[tt]p_memsz[e] against the actual file size and against the one private region every process's address space now reserves for it ([tt]0x400000[e]-[tt]0x2000000[e], the same bounds [tt]vmm_map_page[e] itself enforces, [tt]vmm.nsc[e] -- the upper boundary still matching [tt]alloc.nsc[e]'s own [tt]heap_arena[e] start) before a single byte gets copied anywhere. [tt]elf_window_lo[e]/[tt]elf_window_hi[e] and the old per-program split are gone along with the two fixed windows they used to bound: real isolation now comes from separate physical frames and separate page tables per process (see [tt]vmm.nsc[e], above), not from which half of one shared window a program's segments happened to fall inside, so this check is down to the one region every process shares the same way. that bounds check didn't always exist, and its absence was a real, if never-triggered, gap: [tt]p_memsz[e] costs nothing on disk (unlike [tt]p_filesz[e]), so a program with an oversized [tt].bss[e] could in principle zero straight through [tt]heap_arena[e] from inside that same shared private region -- nothing this project's own toolchain ([tt]nscc[e] plus its own link script) actually produces gets anywhere near triggering it, but it's cheap to close and matches this kernel's usual "reject cleanly rather than trust blindly" convention elsewhere, so it's closed now rather than left as a known theoretical gap. the check itself is written to stay correct against a garbage or malicious 64-bit [tt]p_memsz[e] too, not just a well-formed one: it confirms [tt]p_memsz[e] already fits the window [i]before[e] rounding it up to 8 (what the zero-fill loop below actually needs), so a value near the top of the 64-bit range can't wrap the round-up back around to something small enough to slip past the check.
 
past validation, it walks every [tt]PT_LOAD[e] program header, copies [tt]p_filesz[e] bytes from the file to [tt]p_vaddr[e], and zeroes whatever's left up to [tt]p_memsz[e] (the segment's own [tt].bss[e]) -- starting that zero-fill from [tt]p_filesz[e] rounded [i]up[e] to the next 8-byte boundary, not from [tt]p_filesz[e] directly: an 8-byte-stride loop starting from an unaligned offset could overshoot [tt]p_memsz[e]'s own validated bound by up to 7 bytes on its last write, a real, narrow gap in the bounds check above that a later pass found and closed the same way the copy loop right before it already rounds. no relocations, no dynamic linking, no sections -- a plain static non-PIE executable only, which is also all [tt]elf_call_entry[e] can actually run anyway: there's no ring3/tss/user segments yet, so a "loaded" program runs in ring0 as an ordinary called function ([tt]call *entry[e], not a real process replacement) and is expected to [tt]ret[e] back when it's done, the same simplification as everywhere else in this kernel that doesn't have a scheduler. the byte-level reads ([tt]elf_read_u16[e]/[tt]u32[e]/[tt]u64[e]) exist because elf's layout is dictated by the format spec, not kaboom's own invention (unlike kfs) -- so this needs the same byte-extraction technique as gdt/idt parsing, not kfs's word-aligned shortcuts.
 
[h2]drivers[e]
 
[h3]vga (src/drivers/vga.nsc)[e]
 
text-mode vga at [tt]0xb8000[e], 80x25, 2 bytes per cell (char, attribute). every write is a read-modify-write of the containing 8-byte-aligned word, because nsc's [tt]*ptr[e] is always a full 8-byte load/store -- there's no byte or halfword-granular deref, and structs can't model a packed hardware layout either since nsc struct fields are word-addressed, not byte-packed. cursor positioning and scrolling are a direct port of tape-kernel's own [tt]cm.c[e]/[tt]vga.c[e] ([tt]cnb[e]/[tt]cob[e]/[tt]scur[e]/[tt]hcur[e]/[tt]scrl[e]), not reinvented; the one thing not ported is [tt]cnb[e] (reading the cursor position back from the crtc), since [tt]vga_row[e]/[tt]vga_col[e] already track kaboom's own idea of cursor position and nothing else ever moves the hardware cursor independently.
 
[h3]serial (src/drivers/serial.nsc)[e]
 
com1, port 0x3f8, a 16550 uart. started life as a debug output channel only -- boot.s's own raw beacon byte proves the pipeline works before kmain even runs -- but it's also a real input channel now, and that mattered for a real reason: under [tt]-nographic[e], com1 is the [i]only[e] input channel there is, since there's no separate graphical window with its own ps/2-backed keyboard focus. that gap went unnoticed for a while because every keystroke this project's own testing had ever used came from qemu's QMP send-key interface, which injects real ps/2 scancodes regardless of display mode -- an actual person typing into a real [tt]-nographic[e] session was never exercised until someone tried it and nothing happened at all. [tt]serial_handle_irq[e] (irq4) now translates a carriage return (13) to a newline (10) and both real backspace (8) and delete (127) to the same byte, because a raw-mode host terminal does none of the translation a cooked terminal normally would -- and drops anything outside backspace/tab/newline/printable-ascii before it's ever echoed, specifically because echoing a bare [tt]esc[e] (27, what both arrow keys and delete start with) back to the user's real terminal let them navigate into and "edit" output that had already scrolled past, purely as a trick of their own terminal's rendering; kaboom's own line buffer was never actually touched by any of it. [tt]serial_putc[e] also translates a bare newline to a carriage return followed by a newline on the way out, for the same raw-mode reason in reverse -- found by actually cat-ing a multi-line file over [tt]-nographic[e], not by reading the code, since a captured-to-file serial log (this kernel's only test method for a while) never shows a live terminal's cursor going wrong.
 
[h3]kbd (src/drivers/kbd.nsc)[e]
 
ps/2 keyboard, scancode set 1, us qwerty, make codes only -- break codes are read (so the controller doesn't stall) and mostly discarded, except shift's own release, which has to be tracked. the scancode-to-ascii table, shift, and caps lock are a direct port of tape-kernel's [tt]kb.c[e] ([tt]scntasci[e]/[tt]gtchr[e]), reshaped into an if-chain instead of a switch since this is irq-driven, not [tt]gtchr[e]'s own polling loop -- the actual mapping is identical. extended (0xE0-prefixed) scancodes are still just discarded, matching tape-kernel's own "for now." every mapped keypress goes through one shared function, [tt]kbd_push[e], which both echoes it (vga + serial) and buffers it into a 256-byte circular buffer -- the same function [tt]serial.nsc[e]'s own irq handler calls for a byte that arrived over com1 instead, so neither [tt]kbd_getchar[e] nor its readers ever need to know which physical path a byte actually came in on. backspace is deliberately [i]not[e] echoed here -- that's [tt]fd_read_stdin[e]'s job, once it actually knows whether there's anything on the current line worth erasing, which an irq handler firing on every keystroke has no way to know.
 
[h3]ata (src/drivers/ata.nsc)[e]
 
ata/ide pio, primary bus, master drive, lba28, polling only -- no dma, no irq-driven i/o. a clean-room rewrite of tape-kernel's own approach, not a port. a 512-byte sector is 64 8-byte words; four consecutive 16-bit port reads from [tt]0x1f0[e] get packed into one [tt]u64[e] before a single [tt]*ptr[e] store (and symmetrically unpacked on write), the same packed-word technique vga uses, since there's still no byte/halfword-granular deref available.
 
[h3]rtc (src/drivers/rtc.nsc)[e]
 
cmos real-time clock, ports 0x70 (index) and 0x71 (data) -- the same pair every pc-compatible has had since the original at. polling only, no periodic-interrupt mode. [tt]rtc_read_datetime[e] checks "update in progress" (status register [tt]0x0a[e], bit 7) before reading, which avoids the real failure mode -- torn, nonsensical values from reading mid-update -- but there's one race left deliberately unhandled: reading all six fields isn't atomic, so a read that straddles the clock actually ticking over (say, catching [tt]23:59:59[e] right as it rolls to the next minute, then reading a now-stale day/month/year) can come back very slightly wrong, extremely rarely. that's an accepted imprecision, not worth a read-twice-and-compare loop for what's still a debug-grade clock with nothing syncing to it. year assumes the 2000s, same as every other from-scratch rtc reader that doesn't bother with the century register.
 
[h3]cpu (src/drivers/cpu.nsc)[e]
 
cpu identification via [tt]cpuid[e], vendor and brand strings only -- no feature-flag decoding (leaf 1's edx/ecx bits, leaf 7's ebx/ecx/edx, none of it), matching this kernel's "ship the real minimum" pattern everywhere else. [tt]cpu_vendor[e] reads leaf 0 and has to know its own well-known quirk: the 12-character vendor string comes back in ebx, edx, ecx register order, not the eax/ebx/ecx/edx order every other multi-register cpuid result (including the brand string) actually uses. [tt]cpu_brand[e] reads the three extended leaves [tt]0x80000002[e]-[tt]0x80000004[e] in normal register order for the 48-character brand string, and doesn't bother checking leaf [tt]0x80000000[e]'s own max-supported-leaf return first -- every cpu qemu emulates supports these, and a real cpu old enough not to would already have failed this kernel's boot for unrelated reasons long before cpu identification became the problem. both are exposed as [tt]/int/cpu[e] (virtfs), and [tt]info[e] cats it straight through -- confirmed against a real qemu boot to correctly report the actual host cpu qemu is emulating ([tt]AuthenticAMD[e] / [tt]QEMU Virtual CPU version 2.5+[e], on the machine this was tested on), not a hardcoded string.
 
[img src="made-with-nsc.gif"]made with nsc[e] [img src="powered-by-kaboom.gif"]powered by kaboom[e]
powered by btf.