| git.druid.rocks | index | druid520 | kaboom | docs/ | fs.btft |
docs/fs.btft
[link rel="stylesheet" href="keyframes.css"][e]
[table class="topnav"]
[tr]
[td class="logotab"][see name="index"]kaboom[e][e]
[td][see name="kernel"]kernel[e][e]
[td][see name="fs"]fs[e][e]
[td][see name="syscalls"]syscalls[e][e]
[td][see name="userland"]userland[e][e]
[td][see name="build"]build[e][e]
[e]
[e]
[h1]fs[e]
this page is kfs: kaboom's own on-disk filesystem, hand-rolled and not a port of anything -- not fat12/fat32, not tape-kernel's own ffs (clean-room, feature-inherited-not-copied, since tape-kernel is gplv3 and this is fmc), not unix's inode format even where it looks similar. it also covers virtfs, the small layer that makes [tt]/proc[e] and [tt]/int[e] show live kernel state instead of real disk content. read [see name="kernel"]kernel[e] for how exec.nsc drives a lot of this (path resolution, permission checks) on the way to running a program.
[h2]on-disk layout[e]
block 0 is the superblock, blocks 1 through inode_count are the inode table (one inode per whole 512-byte block -- see below), the next block is the root directory's own data, and everything past that is file data blocks, bump-allocated and never freed. every field in every one of these structures is a full 8-byte word, deliberately: unlike gdt/idt entries, whose byte layout is dictated by the cpu, or elf, whose layout is dictated by the format spec, kfs's on-disk format is entirely kaboom's own invention, so there's no reason to fight nsc's pointer model ([tt]*ptr[e] is always a full 8-byte load/store, no byte or halfword deref, no struct byte-packing) when word-aligning every field sidesteps the problem for free.
an inode is a whole 512-byte block -- not several inodes packed per block the way a real unix filesystem would do it. that costs disk space nobody's short on yet, and buys back something real: reading or writing an inode is a plain, direct block read or write, no sub-block offset math, no read-modify-write of a block shared with unrelated inodes. the fields: type (0 free, 1 file, 2 directory), permission bitmask, size in bytes, 56 direct block pointers, 1 single-indirect pointer, and 4 double-indirect pointers -- 61 pointer slots total, unchanged from the 512-byte inode's own size. that cap has been raised four times now -- 64 bytes/5 direct blocks, then 128/13, then 512/60 (single-indirect: 60 direct + 64 more via one indirect block, 124 blocks = 63488 bytes), each time because kaboom's own coreutils binaries genuinely outgrew the old one, and most recently double indirection: 4 of the 60 direct slots became double-indirect pointers instead, each reaching 64*64 = 4096 blocks (2mib) through a two-level 64-pointer fan-out, for 56 + 64 + 4*4096 = 16504 blocks = 8450048 bytes -- the format's ceiling, not what's actually reachable end to end today (see below). [tt]kfs_block_tier[e] is the one place this tier arithmetic (which of the three regions a given logical block number falls in, and its exact slot within it) is computed at all; [tt]kfs_write_file[e], [tt]kfs_read_file[e], and [tt]kfs_scan_allocators[e] all call it instead of separately re-deriving the same boundaries, so the three can never silently drift out of agreement with each other the way a hand-copied formula in three places eventually would. [tt]mk/disk.pl[e] (a separate language, can't call nsc) mirrors the identical arithmetic in perl for baking files in at build time -- kept in sync by hand, the same "no shared source of truth, same discipline anyway" relationship it already has with the rest of [tt]kfs.nsc[e].
the 8450048-byte number is the on-disk format's own ceiling, not a promise that a file that big can actually exist end to end: a single open file is capped by [tt]fs_arena[e] (4mib, see [see name="kernel"]kernel[e]'s [tt]alloc.nsc[e] coverage), a fresh default disk image only has roughly 1.5mib genuinely free once everything baked-in is accounted for, and one boot session's total budget for large whole-file reads is bounded by the heap's own non-freed remainder past [tt]fs_arena[e]'s carve-out. each of those is independently raiseable later, same "extend when something actually needs it" pattern as the tier format itself -- not attempted here since nothing needs it yet, and it's worth saying plainly rather than letting the headline number oversell what's really usable today.
this was raised for a real reason, not preemptively: a planned real-bios bootloader ("dynamite", see [see name="build"]build[e]) needed to embed a raw kernel blob well past the old 63488-byte ceiling. two real, severe bugs turned up building this, both found by deliberate post-landing review rather than assumed away: [tt]kfs_dir_add_in[e] used to corrupt the directory block outright for any filename over 23 bytes (the null-terminator write used the untruncated length, writing past the truncated name field -- in the worst case, into the directory chain's own "next block" pointer, misreading a real disk region as dirents on the next walk); and neither [tt]kfs_mkdir[e] nor [tt]kfs_dir_add_in[e]'s own chain-block allocation checked against the disk's real size the way [tt]kfs_write_file[e] already did, so a nearly-full disk could link a directory chain block past the image's own end and leave the next boot unable to mount at all. both are fixed now -- an over-length name is refused cleanly before [tt]kfs_dir_add_in[e] is ever reached, and directory block allocation shares the same disk-size bound file writes already had -- but they're documented here rather than quietly folded in, since a filesystem format change is exactly the kind of thing worth being honest about what broke along the way.
[h2]permissions: a bitmask, not unix octal[e]
kfs permissions are a single bitmask digit, r=1 w=2 x=4, so rwx is 7 -- the same scheme for files and directories both, stored directly in the inode, no separate owner/group/other split anywhere. [tt]chmod.nsc[e]'s own comment says the important part plainly: this is [i]not[e] unix's three-digit owner/group/other octal, and treating it like one is a real, documented point of confusion -- kaboom has no user accounts to have separate owner/group/other permissions for in the first place, so one digit is the whole permission model, unconditional, with no owner/root gate to check against first. there's exactly one user here, so the stored bitmask is simply the whole answer, every time.
that distinction bit someone for real, not hypothetically: several places that predate real enforcement -- [tt]mk/disk.pl[e]'s own bin-binary/doc/virtfs-placeholder creation, and [tt]kfs_save[e]'s create-if-missing default -- stored the literal permission "6" believing it meant "rw", which is true in real unix octal (r=4 w=2) but wrong in kaboom's own scheme, where 6 is actually "wx" with r entirely missing. harmless before enforcement existed to actually check it; it would have made every binary on the disk, sh included, and every doc file unreadable and unexecutable the moment enforcement turned on, without anyone asking for that. fixed by correcting the defaults (disk.pl's binaries/docs/placeholders now get 7, kfs_save's create-default now gets 3) rather than by touching the bitmask itself.
[tt]chmod[e] (the userspace command) takes either the raw digit ([tt]chmod 3 file[e]) or the same bitmask spelled out as an rwx letter triplet with [tt]-[e] for an unset bit ([tt]chmod rw- file[e]) -- exactly the form [tt]perms[e] prints back, so whatever [tt]perms[e] shows you is also valid [tt]chmod[e] input, unchanged either way. [tt]perms[e] itself only exists because nothing before it could read a permission bit back out to userspace at all: [tt]kfs_check_perm[e] could check one, [tt]kfs_chmod[e] could overwrite one, but neither could hand one back for a command to display.
[h2]path resolution -- and the gap that's still there[e]
[tt]kfs_dir_find[e]/[tt]kfs_dir_find_in[e] only ever look inside one given directory; they can't walk into a subdirectory themselves. [tt]kfs_resolve[e] is the function that actually walks a real path one component at a time, absolute ([tt]/a/b/c[e], starting at root) or relative ([tt]a/b/c[e] or a bare name, starting at cwd), calling [tt]kfs_dir_find_in[e] once per component. every intermediate component has to already exist and be a directory; the final component can be anything, since a file is exactly what [tt]cat[e]/[tt]open[e] want to find at the end of a path. [tt]kfs_resolve_parent[e] is the same walk, except it stops one component early and hands back the parent directory's own lba and inode number plus the byte offset where the final component starts -- what [tt]mkdir[e]/[tt]create[e]/[tt]rm[e]/[tt]rmdir[e]/[tt]save[e] all actually need, since they're operating a dirent into or out of the parent, not resolving the (possibly not-yet-existing, for create/mkdir) final inode itself.
say the gap plainly: [tt]..[e] is not supported. there's no parent pointer stored anywhere -- not in an inode, not in a dirent -- so there's nothing for [tt]..[e] to resolve against even if the parser recognized it, and today it doesn't even try. this is a real, known limitation, not a secret held back from this page. [tt]kfs_cd[e] works around needing a parent pointer at all by keeping [tt]kfs_cwd_path[e] as its own separately-tracked printable string, updated by hand on every [tt]cd[e] (replaced wholesale on an absolute path, appended to on a relative one) rather than ever being reconstructed by walking parent pointers backward -- because no parent pointer exists to walk. [tt]pwd[e] just reads that string back; it isn't derived from anything else.
descending through an intermediate directory during either resolve function costs a real permission check, not just a type check: the component has to actually be a directory, and it has to have the execute bit set -- real unix's own "traverse" permission -- checked with [tt]kfs_check_perm(inode, 4)[e] before the walk is allowed to continue through it. that check is why a directory's x bit matters independently of its r bit at all; see the enforcement section below for the exact list of what depends on which bit.
[h2]directories are a chain of blocks, not one[e]
a directory's data used to be exactly one 512-byte block: sixteen 32-byte dirent slots, each an 8-byte inode number field plus a 24-byte null-padded name, and that was the whole directory, full stop. it broke for a real reason, not a hypothetical one: [tt]/bin[e] outgrew sixteen entries the moment a few more coreutils got added, and [tt]mkdir: directory (lba N) full[e] -- disk.img.def's own guard for exactly the single-block case -- fired for real.
the fix chains blocks instead of growing the inode: slots 0 through 14 (offsets 0 through 448) are real dirents, one fewer than before, and slot 15 (offset 480) is never a real dirent at all -- its full 8-byte word is either the all-ones sentinel (no next block yet) or the lba of the next block in the chain. [tt]kfs_dir_find_in[e]/[tt]kfs_dir_list_in[e]/[tt]kfs_dir_add_in[e]/[tt]kfs_dir_remove_in[e] all walk this chain, only moving to the next block once the current one is exhausted, and [tt]kfs_dir_add_in[e] only bump-allocates and links on a fresh block once the current last block in the chain is genuinely full -- a directory grows by one block at a time, exactly when it needs to, the same lazy-allocation shape [tt]kfs_write_file[e]'s own indirect block already uses.
growing a directory's inode to hold multiple direct block pointers, the way a file's inode already does, was the other option and got passed over on purpose: chaining needed zero changes to [tt]kfs_mkfs[e], to [tt]mk/disk.pl[e]'s own block format, or to anything that already treated a directory as "an lba", like [tt]kfs_cwd_dir_lba[e]/[tt]kfs_bin_dir_lba[e] -- a freshly formatted block's slot 15 was already the all-ones sentinel every slot starts as, so "no next block yet" was already true of every directory block that existed before chaining did. only the four [tt]dir_*_in[e] functions above needed to actually change.
removing entries doesn't shrink the chain back down, either: [tt]kfs_rmdir[e]'s own "is this directory empty" check has to walk every block in the chain, not just the first, because a directory that once grew a second block and then had every entry in it removed again is still a chain of now-empty blocks, not back to one -- the same bump-allocator-never-frees philosophy as every other allocator in this kernel (kalloc, the fs block/inode allocators), documented rather than fixed.
[h2]permission enforcement: exactly where each bit is checked[e]
permissions are genuinely enforced, not just stored -- [tt]kfs_check_perm(inode_num, bit)[e] is the one real check every enforcement point below shares, and it's the [i]only[e] place in the whole kernel that ever sets [tt]kaboom_errno[e] to 1 (EPERM); every public entry point with a permission check of its own resets [tt]kaboom_errno[e] to 0 at its own start, so a stale value from an earlier, unrelated failure can never leak into a later call's error log line (see [see name="kernel"]kernel[e]'s note on [tt]sys_doerror[e] for what actually reads that value).
for a file:
[list]
[list-item]r gates [tt]fd_open[e] -- opening it for reading at all, which covers [tt]cat[e], [tt]ed[e]'s load, and the source side of [tt]cp[e]/[tt]mv[e].[e]
[list-item]w gates [tt]kfs_save[e]'s overwrite-an-existing-file path.[e]
[list-item]x gates [tt]sys_exec[e] -- and that check applies identically to a shebang script's own interpreter re-exec as it does to the original script or binary; there's no separate, weaker path for "the thing sys_exec found on its own."[e]
[e]
for a directory, it's a genuine three-way split matching real unix semantics, not a kaboom invention:
[list]
[list-item]r gates listing it -- [tt]kfs_dir_list[e]/[tt]kfs_dir_list_path[e], what [tt]ls[e] calls.[e]
[list-item]w gates creating or removing an entry within it -- [tt]kfs_create[e]/[tt]kfs_mkdir[e]/[tt]kfs_rm[e]/[tt]kfs_rmdir[e], all checked against the [i]parent[e] directory's own inode via [tt]kfs_resolve_parent[e]'s dedicated parent-inode output parameter, added specifically because the parent directory's data-block lba alone has no way back to its own inode number, which is what [tt]kfs_check_perm[e] actually needs.[e]
[list-item]x gates traversing through it at all -- every intermediate component of any path ([tt]kfs_resolve[e]/[tt]kfs_resolve_parent[e]), plus the final directory [tt]kfs_cd[e] is actually moving into.[e]
[e]
worth being explicit about a real, already-confirmed edge case here: [tt]chmod 0[e] on a [i]file[e] doesn't stop [tt]ls[e]/[tt]stat[e] from showing it, and that's correct, not a bug -- real unix never gated seeing a file's own dirent entry on that file's own permission bits either, only on x on the directories leading to it. [tt]cat[e]-ing or executing that same chmod-0'd file does correctly fail, and correctly works again after [tt]chmod 7[e] restores it; this was checked directly against a real repro ([tt]chmod 0 /bin/touch[e] then [tt]ls /bin/touch[e]) rather than assumed.
[tt]kfs_chmod[e]/[tt]kfs_getperm[e] themselves are unconditional -- no permission check gates changing or reading back a permission bit, because there's no owner or root concept to gate that behind (the same "one user, no split" reasoning as everywhere else in this section). the one user here can always chmod anything.
[tt]kfs_save[e] (what [tt]cp[e]/[tt]mv[e]/[tt]ed[e]'s [tt]w[e] all go through to actually write a file) checks the resolved target's own type before ever touching it, refusing cleanly if it already exists as something other than a plain file -- a directory, most obviously. that check exists because of a real, live-reproduced bug: without it, [tt]kfs_write_file[e] overwrote whatever inode it was handed regardless of type, so "mv file dir" or "cp file dir" -- the ordinary, natural unix habit of moving something [i]into[e] a directory, typed one path short of the real target -- silently replaced the directory's own dirent block with the file's raw bytes, destroying every entry in it. [tt]kfs_save[e] also rolls back its own work now on a different failure: if it had to [tt]kfs_create[e] a brand-new inode first (the target didn't exist yet) and the subsequent [tt]kfs_write_file[e] then failed -- typically a disk nearly full -- it used to leave that freshly-created, permanently empty inode behind. that empty file was a real landmine: an entirely ordinary later [tt]mv[e] onto the same name would treat it as a legitimate source, silently overwriting a real file with nothing. [tt]kfs_save[e] now removes the just-created inode on that specific failure path instead, restoring the pre-call state exactly (the name simply doesn't exist, same as before it was ever called).
[h2]virtfs: /proc and /int[e]
a handful of names under [tt]/proc[e] and [tt]/int[e] have their content generated live, at read time, instead of coming from real disk blocks. real, empty placeholder files still exist on disk for each of these -- [tt]mk/disk.pl[e] creates them at build time -- purely so [tt]kfs_resolve[e]/[tt]ls[e]/[tt]stat[e] keep working completely unmodified on them. [tt]virtfs_read[e] is what [tt]fd_open[e] calls [i]first[e], before ever falling through to a real [tt]kfs_read_file[e], so opening one of these always returns fresh content instead of whatever empty bytes were on disk at mkfs time. [tt]kfs_proc_dir_lba[e]/[tt]kfs_int_dir_lba[e] are resolved once at mount time, the same way [tt]kfs_bin_dir_lba[e] is, so virtfs can recognize "this name's parent is one of these two directories" without re-walking from root on every single open -- and both fall back to root on an old disk image built before [tt]/proc[e]/[tt]/int[e] existed at all, the same fallback [tt]kfs_bin_dir_lba[e] already uses, guarded so a root-level file that happens to be named e.g. "version" on such an image isn't wrongly treated as virtual.
[tt]/proc/version[e] prints a fixed banner string. [tt]/proc/self[e] reports whoever is currently reading it -- plan9's own convention, not a fixed pid -- which falls out for free: anything able to open and read a file at all is, by definition, a running process, so the deepest currently-live slot in kaboom's own process table (see [see name="kernel"]kernel[e]'s own coverage of it) is always the honest answer, whatever the actual nesting depth happens to be right now. [tt]/proc/ps[e] is the whole process table, one line per depth that's actually populated -- "ps is literal [tt]ls /proc[e]", and [tt]ps.nsc[e] itself is nothing more than a hardcoded [tt]cat[e] of this file. deliberately [i]not[e] here: a real per-pid subdirectory tree ([tt]/proc/1/status[e] and so on) would need virtual [i]directories[e], not just virtual files -- [tt]kfs_resolve[e]/[tt]kfs_stat[e]/[tt]ls[e] would all have to learn to fake a listing, not just have [tt]fd_open[e] call a generator -- a materially bigger change for up to ten processes that are already fully described by one flat [tt]/proc/ps[e] line each.
[tt]/int/mem[e] reports live heap and fs-arena usage via [tt]alloc.nsc[e]'s own accessors, so this doesn't need to know the internal layout of either arena. [tt]/int/kbd[e] reports live shift/caps-lock state. [tt]/int/kfs[e] reformats the same four numbers [tt]kfs_info[e] already hands the [tt]info[e] command as plain text. [tt]/int/cpu[e] reports the vendor and brand strings straight from [tt]cpu.nsc[e]'s cpuid wrappers -- confirmed against a real qemu boot to report the actual host cpu qemu is emulating, not a hardcoded string.
[tt]kfs_rm[e]/[tt]kfs_rmdir[e]/[tt]kfs_create[e]/[tt]kfs_mkdir[e] all refuse to touch an entry whose parent is [tt]/proc[e] or [tt]/int[e], reusing the exact same [tt]kfs_proc_dir_lba[e]/[tt]kfs_int_dir_lba[e] recognition virtfs's own dispatch already relies on. that check exists because of a real, live-reproduced bug: without it, an entirely ordinary [tt]rm /proc/ps[e] or [tt]mv /int/mem x[e] permanently deleted the real, disk-backed placeholder dirent these virtual files need to exist at all -- surviving a reboot, since the removal is a genuine disk write, not something confined to one boot session. [tt]touch[e]/[tt]mkdir[e] get the identical refusal for the same reason in reverse: without it, [tt]touch /proc/foo[e] would create a new placeholder inside a virtual directory that nothing could ever remove again, since the removal path is now protected too.
dispatch itself -- which known parent directory, and which exact name within it -- is a plain if/else chain, checked directly against the literal name each time. that's not laziness: nscc, the compiler, has no function pointers at all. every call site has to name a real function symbol at compile time; there's no way to take a function's address, store it in a variable or a struct field, or call through one. nsc also has no arrays of structs -- only arrays of a single scalar or pointer type -- so a [tt]{name, fn}[e] table iterated in a loop isn't buildable in this language as it stands. this was investigated directly against the compiler's own parser and typechecker, not assumed: a real jump table would need an actual nscc compiler change (indirect-call codegen), a much bigger undertaking than de-hardcoding one dispatch function, and arguably works against the same "small, simple" design direction that would motivate wanting a table in the first place. an if/else chain calling a named function per case is already the most direct, minimal shape this language can express for a lookup like this -- restructuring it without real function pointers would only add ceremony, not reduce hardcoding.
[img src="made-with-nsc.gif"]made with nsc[e] [img src="powered-by-kaboom.gif"]powered by kaboom[e]