| git.druid.rocks | index | druid520 | kaboom | docs/ | userland.btft |
docs/userland.btft
[link rel="stylesheet" href="keyframes.css"][e]
[table class="topnav"]
[tr]
[td class="logotab"][see name="index"]kaboom[e][e]
[td][see name="kernel"]kernel[e][e]
[td][see name="fs"]fs[e][e]
[td][see name="syscalls"]syscalls[e][e]
[td][see name="userland"]userland[e][e]
[td][see name="build"]build[e][e]
[e]
[e]
[h1]userland[e]
this page is everything that runs on top of [tt]exec.nsc[e] handing control to a loaded elf: a from-scratch userspace libc, every coreutil built on it, and [tt]sh[e] -- the shell, the shebang interpreter, and the thing that turns a chmod-7'd text file into a runnable script. read [see name="syscalls"]syscalls[e] for the exact calls all of this rests on, and [see name="kernel"]kernel[e] for how a program actually gets loaded and re-exec'd; this page only covers what happens once one is running.
[h2]a from-scratch libc, opt-in per program[e]
[tt]lib.nsc[e] ([tt]ptr_byte_at[e], [tt]argv_get[e], [tt]u64_to_dec[e]) predates everything else here and is linked into every program unconditionally, always has been -- it's the one thing every userspace program needs regardless of what else it does: reading a byte through a raw [tt]ptr[e] (nsc's [tt]*ptr[e] is always a full 8-byte load/store, never byte-granular), pulling a slot out of argv (itself a raw [tt]ptr[e] to an array of 8-byte pointer slots, so reading one is a plain aligned load, no packed-word trick needed the way byte-at-a-time access requires), and printing an unsigned decimal number. it also still keeps its own [tt]my_strlen[e], but only [tt]stdio.nsc[e]'s own [tt]fputs[e] calls it now -- every real program call site has since moved onto [tt]string.nsc[e]'s [tt]strlen[e] instead (below). [tt]my_strlen[e] stays for that one caller rather than pulling [tt]string.nsc[e]/[tt]mem.nsc[e] into every program that only wants buffered stdio and nothing else -- the same "opt-in, no forced dead weight" reasoning the whole module-linking scheme already runs on.
past that, growing kaboom's own capability while staying 100% nsc -- no fork/exec/wait/pipes/signals, ever, so a real csh-style port was never in scope -- meant writing an actual libc from scratch rather than porting one:
[tt]mem.nsc[e]: [tt]memcpy[e]/[tt]memset[e]/[tt]memmove[e]/[tt]memcmp[e], plus [tt]mem_byte_set[e], the write-side counterpart [tt]lib.nsc[e] never needed (nothing there ever wrote through a raw [tt]ptr[e]). [tt]memmove[e] handles overlap by picking a copy direction based on which end overlaps -- the one thing a plain [tt]memcpy[e] is allowed to get wrong.
[tt]string.nsc[e]: [tt]strlen[e]/[tt]strcpy[e]/[tt]strncpy[e]/[tt]strcat[e]/[tt]strcmp[e]/[tt]strncmp[e]/[tt]strchr[e]/[tt]strrchr[e]. [tt]strlen[e] used to be a real duplicate of [tt]lib.nsc[e]'s older [tt]my_strlen[e] -- every real call site (every coreutil, [tt]sh[e]) has since been migrated onto this one instead. [tt]str_eq[e], the other old duplicate, had zero real callers left anywhere and was deleted outright rather than migrated. [tt]strncpy[e] keeps real strncpy's actual (mis)behavior -- pads with nulls if src is shorter than n, does NOT null-terminate at all if src is n bytes or longer -- kept exactly that surprising because anything porting real code expects it. [tt]strrchr[e] matches ch equal to the terminator byte itself by scanning through the terminator's own position instead of stopping before it, which gets real strrchr's "ch is the terminator" case for free instead of needing a special case.
[tt]ctype.nsc[e]: [tt]is_digit[e]/[tt]is_upper[e]/[tt]is_lower[e]/[tt]is_alpha[e]/[tt]is_alnum[e]/[tt]is_space[e]/[tt]to_upper[e]/[tt]to_lower[e]. plain ascii range checks, no locale, no unicode -- nothing in this kernel or userspace has ever needed either -- and no lookup table, since nsc has no static const array initializers to build one from anyway.
[tt]malloc.nsc[e]: a real explicit-free-list-plus-boundary-tags allocator (first-fit search, splitting on allocation, bidirectional coalescing on free -- the standard textbook design), sitting on top of [tt]sys_alloc[e]/[tt]sys_free[e]. [tt]sys_free[e] really is a no-op kernel-side (see [tt]alloc.nsc[e]'s own note), so this file is what makes [tt]free()[e] actually mean something for a userspace program without changing anything kernel-side. every chunk pulled from [tt]sys_alloc[e] is bracketed with permanently-allocated, zero-payload prologue and epilogue sentinels, so [tt]coalesce()[e] never needs to know a chunk's real bounds -- a sentinel's alloc bit is always set, which stops merging exactly at the edge on its own, and lets multiple chunks ([tt]heap_extend[e] can be called more than once) not need to be contiguous.
[tt]stdio.nsc[e]: buffered [tt]putchar[e]/[tt]puts[e] (one [tt]sys_write[e] per flush instead of one per line, flushing on a newline as well as a full 256-byte buffer, not just the buffer filling), unbuffered [tt]fputs[e] for stderr or anything that doesn't want stdout's buffering, [tt]print_err[e] (below), and [tt]i64_to_dec[e]/[tt]u64_to_hex[e]/[tt]put_udec[e]/[tt]put_idec[e] for number formatting straight to stdout. there's deliberately no [tt]printf[e]: nsc has no varargs at all, no [tt]va_list[e], no [tt]...[e] parameters anywhere in the language, so a real variadic printf isn't something this toolchain can express.
[tt]print_err(fd, prefix, generic_reason)[e] is the userspace half of the kernel's own [tt]kaboom_errno[e]/[tt]sys_doerror[e] error-logging path: it calls [tt]sys_errno()[e] (syscall 23, returns [tt]kaboom_errno[e] directly) and prints "prefix: permission denied" if that's really why the most recent syscall failed, or "prefix: generic_reason" otherwise -- call it right after the failing syscall, before making another one, since [tt]sys_errno[e] only ever reflects the most recent call. every genuine syscall-failure message across cat/mv/ls/mkdir/touch/rm/rmdir/ps/chmod/cp/sh/ed is migrated onto it now; input-validation errors that never reach a syscall at all (chmod's own "perm must be 0-7, or rwx letters") are deliberately not, and neither is [tt]sh[e]'s "command not found" (below) -- there's a real, documented reason that one stays separate.
the opt-in linking scheme is [tt]mk/bu.sh[e]'s own decision, explained in its header comment: [tt]lib.o[e] is linked into every program unconditionally, always has been, but the newer modules (mem/string/ctype/malloc/stdio) are opt-in PER PROGRAM based on which of their own [tt].nsh[e] headers a program actually includes. linking all of them into everything unconditionally was tried first and measured to add roughly 19kb of dead code to every single binary -- which pushed [tt]ed[e] over [tt]disk.pl[e]'s binary size cap for zero benefit to programs that never call any of it. [tt]mk/bu.sh[e] greps each program's own source for [tt]"mod.nsh"[e] to decide what it wants, plus one level of dependency resolution on top: [tt]string.nsc[e] and [tt]stdio.nsc[e] both call [tt]mem.nsc[e]'s functions internally, so their object needs [tt]mem.o[e] on the link line even for a program that never includes [tt]mem.nsh[e] directly itself.
that migration also caught a few real bugs along the way: [tt]info.nsc[e]'s [tt]print_stat[e] had all four of its labels declared one byte too long (a hand-counted [tt]sys_write[e] length, now gone since [tt]fputs[e] computes it), [tt]ed.nsc[e]'s own [tt]lines_putc[e] was a byte-for-byte duplicate of [tt]mem_byte_set[e] (now gone, replaced by the shared one), and [tt]ls.nsc[e]'s hand-rolled 24-byte-capped strlen is gone too -- [tt]puts[e]'s own byte-at-a-time scan is safe on a kfs dirent name without it, since [tt]kfs_name_matches[e] kernel-side already refuses any name over 23 bytes, guaranteeing the null terminator sits well inside the 24-byte field.
[h2]coreutils[e]
most of these are deliberately small -- one argument, no flags, "ship the real minimum" -- so a few genuinely are just a hardcoded read of a virtual file dressed up as a command:
[tt]cat[e] takes one filename (bare, relative, or absolute -- [tt]kfs_resolve[e], by way of [tt]fd_open[e] kernel-side, handles all three identically) and reads it whole into a [tt]sys_alloc[e]'d buffer via [tt]read_all[e] (below), sized to the real file rather than a fixed cap. that wasn't always true, and the fix closed a genuinely severe bug: [tt]cat[e]/[tt]cp[e]/[tt]mv[e]/[tt]sh[e] all used to read at most 63488 bytes in one [tt]sys_read[e] -- fine while kfs itself couldn't produce a bigger file, but once double indirection raised that ceiling (see [see name="fs"]fs[e]), the same cap silently truncated anything larger. [tt]cat[e] just showed truncated output; [tt]cp[e] made a truncated copy; [tt]mv[e] was the severe one -- it wrote the truncated copy, then deleted the only full original, real permanent data loss. the fix: a new syscall, [tt]fsize(fd)[e] (returns an open file's real size), plus [tt]read_all(fd, &len)[e] in [tt]stdio.nsc[e], shared by all four -- allocates exactly the real size (never a blanket worst-case buffer, which would exhaust [tt]fs_arena[e] on one big open) and returns a clean failure rather than a partial buffer on any read error. the buffer is heap-allocated rather than a stack local or a global array for the same reason it always was: kaboom's single shared 64kib stack (see [see name="kernel"]kernel[e]) has no room for a buffer this size in any one frame, and a fixed global array that size would sit in every [tt]cat[e] process's own .bss whether or not it's ever actually used -- exactly the case the arena allocator exists to avoid.
[tt]fd_open[e] (kernel-side) also refuses to open a directory for reading at all now -- a second, related bug the size fix surfaced: opening a directory used to "succeed" with zero bytes of content, which meant [tt]mv dir dest[e]/[tt]cp dir dest[e] silently wiped [tt]dest[e] (zero real bytes written over whatever was there) instead of failing. refusing at [tt]fd_open[e] itself closes it for every command that opens a file at all, in one place, rather than each of [tt]cat[e]/[tt]cp[e]/[tt]mv[e] needing its own type check.
[tt]ls[e] with no argument lists the cwd, same as it always has. with one argument it resolves that argument as a real path ([tt]sys_stat[e]/[tt]sys_listdir_path[e], both routed through [tt]kfs_resolve[e], so a bare name, a relative path, and an absolute one all work) and either prints the name back (if it's a file, matching real ls's own behavior on a file argument) or lists it (if it's a directory) -- "ls doc" on a subdirectory of cwd used to silently do nothing useful at all, since that argument handling never existed before. the local names buffer is sized for 64 entries at 24 bytes each, not the older 16 -- directories can now chain past 15 entries into a second block (see [see name="fs"]fs[e]), so 16 stopped being a real ceiling the moment [tt]/bin[e] itself grew past it.
[tt]cp src dst[e] reads src whole (via [tt]read_all[e], above) and writes it to dst via [tt]sys_writefile[e], which creates dst if it doesn't exist. [tt]mv src dst[e] is genuinely copy-then-delete under the hood, not a real rename: kfs has no rename-in-place primitive at the directory-entry level, and no directory-move primitive either. because the source is destroyed at the end, the order is the whole point: read all of src, write all of it to dst and check the write actually succeeded, only then remove src -- any failure before the removal step leaves src completely untouched. before checking [tt]src[e]/[tt]dst[e] for identity (next), [tt]mv[e] first checks whether they're literally the same file via [tt]sys_ino(path)[e] (a new syscall returning the inode number a path resolves to) and does a true no-op if so -- no read, no write, no remove at all. that check replaced an earlier "repair" approach (write, then check dst survived the remove and put it back from memory if not) that could itself lose the file if the disk was nearly full when the repair needed fresh blocks; checking identity up front instead of repairing after the fact closes the whole failure class at its root rather than patching around it. functionally a move -- src is gone afterward, dst has its content -- just not atomic and not O(1) the way a real rename would be. good enough until something actually needs better.
[tt]mkdir name[e] takes no [tt]-p[e], no mode argument, and always creates with rwx (7) -- [tt]chmod[e] after the fact if something more restrictive is actually wanted. [tt]touch name[e] creates the file if it doesn't exist and is a no-op, not an error, if it does; kfs has no mtime field on an inode at all yet, so unlike real touch this can't update an existing file's timestamp -- "make sure this file exists" is the whole feature for now. [tt]rm name[e] only removes files ([tt]kfs_rm[e] verifies the inode's own type is a file); [tt]rmdir name[e] is the directory counterpart, and relies on [tt]kfs_rmdir[e] already refusing a non-empty directory rather than checking that itself.
[tt]chmod[e] and [tt]perms[e] are a matched read/write pair for the same permission bits. kaboom's permission model is one bitmask digit, 0 through 7 (r=1 w=2 x=4) -- not unix's three-digit owner/group/other octal, since kaboom has no user accounts to have separate owner/group/other permissions for in the first place. [tt]chmod perm file[e] takes that digit directly ("chmod 3 file"), or the same bitmask spelled out as an rwx letter triplet with a dash for an unset bit ("chmod rw- file"), auto-detected by argument length. [tt]perms file[e] is the read-back half: before it existed, nothing could ever read a permission bit back out to userspace at all -- [tt]kfs_check_perm[e] only ever checked one internally, and [tt]kfs_chmod[e] could only write one -- so there was no way to see what a previous chmod actually left a file at, short of a later operation failing or succeeding. [tt]perms[e] prints both forms at once ("bin/sh: 7 rwx"), which means whatever [tt]perms[e] shows you is also valid [tt]chmod[e] input, unchanged either way. building this caught a real, live bug: mixing buffered [tt]putchar[e] (only flushes on a newline, a full 256-byte buffer, or an explicit [tt]stdio_flush[e]) with unbuffered [tt]fputs[e]/[tt]sys_write[e] on the same output line reorders the output, since an immediate write can leapfrog a still-buffered byte sitting ahead of it in program order -- [tt]perms[e]'s first attempt printed "qtest: rw-3" instead of "qtest: 3 rw-" because the buffered digit got pushed all the way to the end. fixed by keeping the whole line on one discipline throughout (fputs/sys_write only, no putchar at all) -- worth knowing for anything else that ever builds a line out of more than one write: pick buffered-throughout or unbuffered-throughout, never interleave the two on one line.
[tt]ps[e] is "literal [tt]ls /proc[e], essentially" -- more literal than that description even suggests: [tt]/proc/ps[e] is a real virtual file, regenerated fresh on every read straight from [tt]exec.nsc[e]'s own process table, and [tt]ps.nsc[e] is nothing but a hardcoded open-read-write of it. [tt]pwd[e] is a two-line wrapper around [tt]sys_pwd[e]. [tt]date[e] prints "YYYY-MM-DD HH:MM:SS" straight off the cmos rtc ([tt]sys_date[e], syscall 22, fills a 6-[tt]u64[e] buffer with second/minute/hour/day/month/year) -- one fixed, sortable format, no formatting options and no timezone handling at all, since the rtc itself is usually just whatever the host gave it (utc under qemu, most of the time). [tt]info[e] dumps kfs stats (blocks and inodes, total and used) via [tt]sys_info[e], then cats [tt]/int/mem[e] and [tt]/int/cpu[e] straight through -- again, showing memory and cpu info is nothing more than reading two virtual files that already format themselves as plain text, the same reuse [tt]ps[e] already relies on. [tt]klogs[e] is dmesg-style: dumps the entire kernel log in one shot, no follow mode. [tt]clear[e] wraps [tt]sys_vga_clear[e] (syscall 16) -- that syscall existed before [tt]clear[e] did, because [tt]ed[e] needed it first to redraw its own screen, but had no standalone command wrapping it until asked for directly. [tt]echo[e] joins its arguments with single spaces and a trailing newline; no [tt]-n[e] flag.
[tt]ed[e] is a real line editor, a genuine port -- not a reinvention -- of tape-kernel's own line editor (its [tt]src/usr/editor.c[e]). the real logic is unchanged: a fixed 32-line by 80-column buffer, [tt]e[e] re-enters every line fresh (an empty line stops early), [tt]w[e] joins the buffer back with newlines and saves it, [tt]q[e] quits. kaboom has no arrays-of-arrays (array elements have to be a scalar type), so the 32x80 buffer tape-kernel declared as a real two-dimensional char array is one flat [tt]i8[e] buffer here, sized 32 times 81, indexed by hand as line-times-81-plus-column -- same data, same layout, just addressed the way every other flat table in this kernel already is. two small, real additions past the original port: [tt]a[e] appends lines after whatever's already there instead of always starting over, and the display now shows each line's own number -- both asked for directly, not part of the port. one honest adaptation, not a logic change: tape-kernel's version repaints the whole screen at fixed row/column positions every loop and reads single keys with a raw, no-echo keyboard read. kaboom's vga driver only exposes sequential, scrolling output plus a hardware clear -- there's no absolute-position write or a real cursor api yet -- so [tt]ed[e] prints the same information (filename, hotkeys, current lines) sequentially after a clear instead of redrawing it in place; a single-byte [tt]sys_read[e] already blocks for exactly one keystroke, which is exactly what the original's own read gave it. building the append feature surfaced a real bug: an earlier version used the loop index hitting 32 as its own stop signal and then returned that same index as the new line count, so stopping normally via an empty line -- not by filling all 32 slots -- always reported a line count of 32 instead of how many lines were actually entered. every save afterward appended one extra newline per phantom line past the real content, and it was found by actually cat-ing a saved file and seeing roughly 29 trailing blank lines, not by reading the code. every working buffer of real size here (the line table, the load buffer, the save buffer) is [tt]sys_alloc[e]'d now, not a stack local or a fixed global array -- an earlier version did exactly that anyway, before the allocator was actually wired up to userspace.
[tt]ed[e] refuses to open a file at all -- not "opens it, silently drops what doesn't fit" -- if its real size exceeds the 32-line/79-column/4096-byte capacity, if it has more than 32 real lines, if any line is over 79 columns, or if it contains a nul byte. that refusal replaced a real, live-reproduced data-loss bug: an earlier version silently truncated anything past its own capacity on load with no warning at all, and [tt]w[e] then overwrote the [i]original[e] file with that truncated fragment -- [tt]cp /bin/ls lscopy; ed lscopy; w[e] left [tt]lscopy[e] a few bytes long, destroying the real 16-kilobyte-plus binary. a second, related bug got fixed the same round: [tt]ed[e] used to treat [i]any[e] failure to open an existing file -- permission denied, and before this fix, oversized too -- identically to "doesn't exist yet," starting a blank buffer that [tt]w[e] would then happily save over real content. only a path that genuinely doesn't resolve at all still starts a blank buffer now; every other open failure refuses cleanly instead.
[h2]4c: forthc's language, ported[e]
[tt]4c[e] is [tt]forthc[e] (a separate, sibling project -- a self-hosted forth that jit-compiles straight to raw x86-64 machine code and writes its own elf, its own hand-rolled rex/modrm encoder and linux/netbsd syscall abi included) ported to kaboom's userland -- not forthc's code generator, which has nothing to port: nsc can't emit machine code at all, and kaboom already has [tt]nscc[e] for producing real binaries and [tt]elf.nsc[e] for loading them (see [see name="kernel"]kernel[e]). what's actually ported is forthc's [i]language[e] and its compile-time structure: tokenize, then compile with the same control-stack backpatching forthc's own [tt]do_if[e]/[tt]do_else[e]/[tt]do_begin[e]/[tt]do_do[e]/[tt]do_colon[e]/... use, and the same compile-time-resolved [tt]do[e]/[tt]loop[e] nesting depth -- quirks included, not smoothed over: a recursive word's own [tt]do[e]/[tt]loop[e] shares one depth-indexed slot across every recursive level, exactly like forthc's own. the compiled output is a flat array of cells [tt]execute()[e] interprets in the same process, not machine code -- usage is [tt]4c script.fs[e], a batch run of one script, not a repl.
every word [tt]4c[e] knows -- forthc's own primitive stack words (arithmetic, stack shuffling, comparisons) plus its compile-time control words (: ; if else then begin until while repeat do loop i recurse variable) -- is one entry in a single runtime-built function-pointer table: [tt]&f[e] isn't a constant expression in nsc, so the table is filled in at startup rather than statically initialized, same as [tt]test/t53_funcptr.nsc[e]'s own proof-of-concept over in [tt]nscc[e]. this is the reason [tt]nscc[e] grew real indirect-call codegen at all -- [tt]sh.nsc[e]'s own three-way builtin dispatch (above) was checked directly against exactly this question a while before [tt]4c[e] existed and found wanting for it, back when nsc genuinely had no way to take a function's address at all; [tt]4c[e] is what that feature was added for.
kaboom's own execution model forced two real, deliberate departures from forthc, both where forthc would fault instead of behaving: stack under/overflow and division errors print "err: ..." to fd 2 and stop, rather than crashing outright. [tt]bye[e] can't make an exit syscall the way forthc's own does -- kaboom has none; returning from [tt]main[e] is exiting (see [tt]crt.s[e]) -- so it halts the run loop instead and [tt]main[e] returns 0.
[tt]4c[e] doesn't include [tt]stdio.nsh[e]/[tt]string.nsh[e]/[tt]mem.nsh[e] the way most of the coreutils above do: linking all three (roughly 12kb) would have pushed it well past [see name="build"]build[e]'s own binary size cap on its own, before [tt]4c[e]'s own code -- it has its own small buffered-output, decimal-print, and whole-file-read helpers instead, the same idea as [tt]read_all[e] above but sized to fit. [tt]nscc[e] gives every intermediate value its own stack slot, never reused, and every call argument its own slot too, even a fixed 0 (see [tt]nscc[e]'s own [tt]jalloc7[e]) -- real, measured consequences of that, found and fixed by actually comparing compiled sizes rather than guessing: narrower [tt]emit1[e]/[tt]emit2[e]/[tt]cpush2[e]/[tt]cexpect1[e] wrappers around [tt]emit[e]/[tt]cpush[e]/[tt]cexpect[e]'s own common fixed-trailing-argument calls, and two branches that only differed in which constant they emitted merged into one call fed by a variable instead. splitting a long function into smaller ones and hoisting a repeated global read into a local were both tried too, and both measured [i]worse[e], not better -- left un-done on purpose once that was known, not missed. even after real trimming, [tt]4c[e] still didn't fit the original 30720-byte cap; see [see name="build"]build[e] for the (small, generic) bump that gave it room.
every word/opcode index [tt]4c[e] uses -- forthc's own primitives, [tt]4c[e]'s internal cell ops, its dictionary-entry kinds -- is a real nsc [tt]enum[e], the first file in this kernel to use one: turned out [tt]struct[e]/[tt]typedef[e]/[tt]enum[e]/[tt]const[e] are all genuinely implemented in [tt]nscc[e] already, confirmed directly rather than assumed -- [tt]nscc[e]'s own docs said otherwise (reserved words with no feature behind them yet) until this was found and they were corrected the same round; true once, just stale by the time [tt]4c[e] was written, unrelated to [tt]sh.nsc[e]'s own dispatch-table note above (that one really is still true today -- [tt]nscc[e] gained real function-pointer codegen specifically for [tt]4c[e], see above, but genuinely had none at all when [tt]sh.nsc[e]'s three-way builtin check was last looked at). everywhere a fixed-size buffer is needed instead ([tt]enum[e] values fold to real compile-time constants, but nsc still requires an actual literal decimal token for an array size, not a named one), the array declaration itself repeats the number as a bare literal, same as every other magic number in this kernel gets named -- a comment, not a symbol.
[h2]sh: builtins, exec, and shebang scripts[e]
[tt]sh[e] is "minimalized csh+ash": one builtin set ([tt]cd[e]/[tt]pwd[e]/[tt]exit[e]), everything else exec'd through kfs's own namespace and the cwd model [tt]kfs_cd[e]/[tt]kfs_pwd[e] already track kernel-side. no quoting, no escaping, no pipes or redirection, no job control -- splitting a line on bare spaces only is genuinely the whole parser, same "ship the real minimum" pattern as every coreutil above.
[tt]run_line[e] is the one function both the interactive prompt and script mode actually dispatch through, so a script behaves identically to typing the same lines by hand -- there's no separate "script dialect" anywhere in this kernel. it splits a null-terminated line into up to 16 words in place, then checks the first word against [tt]exit[e] (returns a stop signal to the caller), [tt]cd[e] ([tt]sys_cd[e] with no argument goes to [tt]/[e], with one it goes there or prints "no such directory"), and [tt]pwd[e] (calls [tt]sys_pwd[e], writes the result, flushes) -- anything else falls through to [tt]sys_exec[e].
the line buffer itself is a fixed 256 bytes, and a real, live-reproduced bug used to lurk past that boundary: input beyond 256 characters wasn't rejected or bounded, it just kept reading -- so whatever came after the buffer filled got silently interpreted as a [i]separate[e] command, not an error and not a truncation. a 251-character [tt]echo[e] argument immediately followed by [tt]ls /bin[e], with no real line break typed in between, actually ran [tt]ls /bin[e] as its own command -- exactly the shape that could just as easily run something destructive the user never actually typed as a separate line. fixed by rejecting the whole line outright once it's too long ("sh: line too long, ignored") and discarding everything up to the real newline, so the excess can never resurface as an unintended second command; a small, genuine off-by-one write past the buffer's own end got fixed in the same round.
that three-way builtin check is a hardcoded [tt]strcmp[e] if/else chain, and this was checked directly, not assumed to be fine: could it be a real dispatch table instead? no -- nscc, the compiler, has no function pointers at all. every call site has to name a real function symbol at compile time; there's no way to take a function's address, store it in a variable or a struct field, or call through one. nsc also has no arrays of structs, only arrays of a single scalar or pointer type, so a name-to-function table iterated in a loop isn't buildable in this language as it stands -- the same finding [see name="fs"]fs[e]'s own virtfs dispatch coverage already documents. a real jump table here would need an actual nscc compiler change (indirect-call codegen), a much bigger undertaking than de-hardcoding one three-branch dispatch, and arguably works against the same "small, simple" design direction that would motivate wanting a table in the first place. an if/else chain calling a named function per case is already the most direct, minimal shape this language can express for either dispatch -- not laziness, the correct choice given what nsc actually is.
when none of the three builtins match, [tt]run_line[e] calls [tt]sys_exec[e] with the first word, its length, the word count, and the word array itself. if that comes back -1, [tt]sh[e] deliberately does NOT route it through [tt]print_err[e]/[tt]sys_errno[e] the way every other failure message in this kernel now does -- per [tt]sh.nsc[e]'s own comment, exec's -1 return is already documented as ambiguous ([see name="syscalls"]syscalls[e]'s own entry for syscall 7 says the same thing): a real exec failure and a child program that legitimately returned -1 as its own exit code are indistinguishable from this return value alone, so [tt]sys_errno[e]'s value at that point can't be trusted to mean "why sys_exec itself failed" either -- a child that ran fine and made its own syscalls would have left [tt]sys_errno[e] at whatever ITS last syscall set it to, not at whatever [tt]sys_exec[e] itself hit. so [tt]sh[e] just prints a flat "sh: command not found" instead, regardless of which of the two actually happened.
[tt]run_script[e] is what a shebang script -- or "sh scriptname" typed directly -- actually runs under: read the whole file into a [tt]sys_alloc[e]'d buffer in one [tt]sys_read[e], then walk it one line at a time, skipping blank lines and anything starting with a hash (including the shebang line itself, if present), running everything else through [tt]run_line[e] exactly as if it had been typed interactively. a trailing line with no final newline still runs. no pipes, no redirection, no control flow at all -- "super simple, just commands" is genuinely the entire feature, not a summary of a bigger one.
there are two distinct ways a script actually reaches [tt]run_script[e]. the first: the kernel's own [tt]sys_exec[e] ([tt]exec.nsc[e], see [see name="kernel"]kernel[e]) notices a resolved file isn't a valid elf but starts with a hash-bang, reads the interpreter path off the first line -- no shebang-line arguments are supported, so something like a real [tt]/bin/sh -x[e] would just fail, "super simple, just commands" again -- and re-execs that interpreter (in practice, always [tt]sh[e], kaboom's only one) with the word array shifted the same way a real unix execve does: the interpreter first, then the script's own path, then whatever arguments the original call had past its own first argument slot. that's exactly the argc-two-or-more case [tt]main[e] checks for, and it calls [tt]run_script[e] directly. the second way: typing "sh scriptname" at the prompt reaches the identical [tt]main[e] code path a completely different way -- [tt]sh[e] is a perfectly ordinary elf, so [tt]sys_exec[e] loads and runs it completely normally, and [tt]sh[e] itself sees the same argc-two-or-more case once it starts running. a conventional shebang line like "#!/bin/sh" -- an absolute, multi-component path -- resolves correctly because [tt]sys_exec[e]'s path resolution goes through [tt]kfs_resolve[e], not a bare-name-only lookup; when that shebang interpreter is [tt]sh[e] itself, re-executing it into its own fixed load window while the outer [tt]sh[e] is still alive on the call stack is exactly the case [see name="kernel"]kernel[e]'s paging coverage exists for.
that protection didn't always cover both paths equally, and it was a real gap, not a hypothetical one: the remap used to live only in the shebang-specific branch, so "sh scriptname" typed directly loaded a second copy of sh over the running one's own window with no remap at all -- it happened to work only because the freshly-loaded copy was byte-identical to the one already running, not because the two were genuinely kept apart. the fix moved the remap out of the shebang branch and into [tt]sys_exec[e]'s general elf-load path, so any load into sh's window while a copy is already resident there gets a fresh frame regardless of which of the two ways it was reached -- shebang scripts still cost exactly one frame per nesting level, same as before.
[img src="made-with-nsc.gif"]made with nsc[e] [img src="powered-by-kaboom.gif"]powered by kaboom[e]