| git.druid.rocks | index | druid520 | nscc | docs/ | nsc.btft |
docs/nsc.btft
[table class="topnav"]
[tr]
[td class="logotab"]nscc[e]
[td][see name="index"]index[e][e]
[td][see name="nsc"]nsc[e][e]
[td][see name="ir"]ir[e][e]
[td][see name="nscc"]nscc[e][e]
[e]
[e]
[h1]nsc: the language standard[e]
nsc is c89, minus everything implicit, plus a rule for everything the original leaves to the implementation. this page is the standard: when the compiler and this page disagree, this page wins and the compiler is wrong.
[h2]types[e]
the type set: the eight scalar ints, the one untyped pointer, void, and two aggregate forms.
[ul]
[li]i8, i16, i32, i64 - signed integers, 8/16/32/64 bits. no char/short/int/long names, no bool.[e]
[li]u8, u16, u32, u64 - unsigned integers, same widths. (unsigned is a keyword that stands for nothing: it errors with a pointer to these.)[e]
[li]ptr - ONE untyped 8-byte address. there is no pointed-to type, no * constructor, and void* does not exist: ptr is the only spelling, and jtype3 canonicalizes everything to it. a function's address is a ptr too (see pointers).[e]
[li]void - no value, ever.[e]
[li]struct NAME - a struct type (see structs below); the IR spelling is struct.NAME.[e]
[li][N]T - a fixed-size array of N elements of scalar type T (see arrays below); the IR spelling is [N]T.[e]
[e]
[h2]typedef, enum, const[e]
typedef NAME is toplevel-only and follows its alias chain at jscope2 (typedef i32 A; typedef A B; works; a cycle is an error; redefining a name is an error). enum NAME { A, B = expr, C } is toplevel-only: NAME becomes a typedef for i32, the enumerators become i32 compile-time constants folded by jtype3 (an enumerator without = is the previous one + 1, the first is 0; values are constant expressions, may reference earlier enumerators, and a cycle is an error). const qualifies a decl/param/gvar: the name can be read but never assigned or ++/-- (checked at jtype3, direct assignment only -- a const's address still yields a plain ptr, per nsc's untyped pointers). const cannot qualify a function.
[h2]structs[e]
struct NAME { T f; T f; ... }; is toplevel-only, fields are SCALARS (i8..u64 or ptr -- no nested structs, no arrays as fields), and every field occupies exactly one 8-byte word: storage is word-addressed end to end, and the words run DOWNWARD (field 0 at the symbol's address, field k at address - 8*k -- same as stack slots). s.f reads/writes field k at its own width; p->f is (*p).f sugar; &s.f and &a[i] are ptrs. structs assign (word copy) only to the SAME struct type; there is no struct equality, no struct arithmetic, no struct casts, and an initializer list must have exactly one element per field. a struct-returning function may return at most 2 words (16 bytes) -- bigger returns are an error, pass a ptr out-param; params count their words toward the 6-word call limit. a struct-returning call is not a value: it must be assigned directly to a local (or be a local's initializer) -- the result lands straight in the target's slots. a field name used THROUGH a ptr (p->x) must be unambiguous across all the unit's structs, since ptr has no pointee type to disambiguate. sizeof(struct NAME) is 8*nfields.
[h2]arrays[e]
i32 a[N] is a fixed-size array (N positive decimal, toplevel or local, N elements of a SCALAR type -- no arrays of structs, no multi-dim arrays). elements are BYTE-PACKED at their real width (i8=1, i16=2, i32=4, i64/ptr=8) and run UPWARD from the array's base -- so i8 buf[64] is a real 64-byte contiguous buffer, buf[i] is the i-th byte, &buf[i] - &buf[0] is i*width, and sizeof(a) is N*width. a[i] is a typed lvalue of the element type; a CONSTANT index outside [0, N) is a compile error, a dynamic index is the programmer's business. arrays are not values: a bare a in an expression is an error (use a[i] or &a[0]); there is no decay, no whole-array assignment or parameter passing. initializer lists must have exactly N elements.
[h2]void rules, in full[e]
void is valid as a function's return type and as the sole member of a parameter list ((void) = zero params), and nowhere else. never a variable type, never a value, never an expression, never sizeof's operand -- all of those are compile errors. a parameter list is explicit: () is an error, zero params is written (void); a void among other params is an error too.
[h2]keywords and identifiers[e]
the keyword list is the single source of truth in jlex0.pl; this page enumerates it:
[code]
i8 i16 i32 i64 ptr void
global if else while return break continue switch case default
struct typedef unsigned enum const sizeof include
[e]
struct, typedef, enum and const are all real (see typedef/enum/const and structs, above) -- this sentence used to list them as reserved-with-no-feature-yet, back before any of the four existed; only unsigned is still genuinely that: it lexes as a keyword so it can never be an identifier, but using it as a type is a clear, deliberate error ("unsigned types are spelled u8/u16/u32/u64") rather than silently accepted. sizeof and include are both real too (see below and the include section). everything else is a keyword of the language proper.
identifiers are [A-Za-z_][A-Za-z0-9_]*. a leading single _ is fine; a leading double underscore is reserved for the implementation and is an error.
there are no trigraphs (explicitly), and no line splicing: a backslash never joins lines -- inside a string literal it begins an escape sequence and nothing else. comments are /* */ only: no //, and no nesting (the first */ closes the comment). an unclosed comment is an error.
[h2]include[e]
include "path.nsh"; -- alone on its own line, nothing else -- splices the full, recursively-expanded content of path in at that point, textually, before real lexing ever starts (jlex0's own preproc step does it). path resolves relative to the file containing the include line, not the compiler's working directory: a header can include a sibling header by its own bare name no matter where nscc was actually invoked from, the same rule c's #include "..." (not <...>) has always used. this is a real reserved keyword (not a `#`-prefixed preprocessor line) specifically so a malformed include directive -- a missing semicolon, an unterminated quote, anything that doesn't match the exact required shape -- surfaces as a clear parse error instead of nscc silently trying to parse the word "include" as some other kind of statement.
the usual use: put a group of bodyless prototypes (the extern-style declarations described under global variables and functions and calls) in a .nsh file, and include it from both the file that calls those functions and the file that defines them. the defining file's own global keyword on the real definition is what still controls export/storage, exactly as if the prototypes had been retyped by hand -- textual inclusion, nothing more.
three things to know going in:
include is the entire preprocessor. there is no #define, no function-like or object-like macro, no token substitution, no conditional compilation (#if/#ifdef), no #line, nothing else c's preprocessor does -- include's one job is textual splicing of a named file, full stop. a name that needs to mean the same literal value everywhere is a real initialized global, not a macro (const is reserved, see keywords and identifiers, so it isn't that either yet); a block of code that needs to run conditionally is a real if, checked at runtime, not compiled away by a #ifdef.
repeated includes of the same file are NOT deduplicated -- including the same .nsh twice (directly, or once directly and once through two different headers that both include it) redeclares everything in it twice in the same translation unit, which is a real redeclaration error unless the two declarations happen to agree exactly (jscope2 already tolerates a repeated, consistent prototype -- see global variables and functions and calls -- so a diamond include of pure declarations is fine; anything with a body or an initializer is not, same as writing it twice by hand would be). there is no include-guard mechanism (no #ifndef/#pragma once equivalent) -- keep .nsh files to declarations only and this never comes up in practice.
error positions (file:line:col) reported by every stage after jlex0 are positions in the fully-expanded text, not the original line inside whichever .nsh the error is actually in -- the same rough edge c's own preprocessor had before #line markers existed. worth knowing when a reported line number looks wrong for the file you're actually looking at.
[h2]literals[e]
[ul]
[li]integer literals: decimal or 0x-hex, no suffixes, no floats. the typing rule: a literal that fits in i32 is i32, otherwise it must fit in i64 (a literal is always non-negative -- negation is a separate unary op, never part of the token -- so "fits in i64" means fits in i64's positive range, up through the classic 0x7fff... ceiling) and is i64, otherwise it is an error. an i64 literal cast down to a genuinely narrower int type is an error -- (i32)5000000000 does not compile; the wrap it would silently perform is exactly the kind of accident nsc exists to catch. (small i32 literals may cast down freely: (i8)200 is deliberate and compiles.) u64 is exempt from this, same as i64 itself: it's the same width, not narrower, so (u64) of any i64 literal -- including one with the high bit of the word set, like (u64)0x8000000000000000 -- is a same-width reinterpretation, never a truncation, and always compiles.[e]
[li]char literals: exactly one (possibly escaped) char; escapes are \n \t \r \0 \\ \' \" \xNN. type i32, like c89's int-promoted 'a'.[e]
[li]string literals: "..." with the same escapes, NUL-terminated, type ptr. address stability: two occurrences of the SAME literal content are the SAME address (jlower5 reuses one .LCn blob); distinct contents are distinct blobs -- no suffix merging.[e]
[e]
[h2]conversions[e]
every conversion that happens anywhere is a literal (cast T ...) node in the tast, inserted by jtype3 -- nothing converts invisibly in a later stage. the implicit ones jtype3 inserts for you: assignments, decl inits, returns, call args, operand unification, condition contexts. three conversions are NEVER implicit, and writing them without the cast is an error: int to ptr (see null pointer below), i64 literal to a narrower int (above), and signed to/from unsigned in a binding or a compound assign -- u8 x = 10; and i8 y = (i8)x * 2; are fine, but i8 x; u8 y = x; needs the cast (a literal adopts the target's signedness, so u8 x = 200; works, while i8 x = -1; needs (i8)-1). casts between int widths truncate, sign- or zero-extend, and reinterpret the sign bit; casts between ptr and i64 are free.
[h2]the null pointer[e]
nsc has exactly one null pointer: the explicit cast (ptr)0. implicit int-to-ptr conversions do not exist: ptr p = 0;, p == 0, return 0 from a ptr function, and passing 0 for a ptr parameter are all compile errors -- write (ptr)0. (p == p and p != p are fine, as is (i64)p == (i64)(ptr)0.)
[h2]arithmetic[e]
the width theorem (a theorem, not an example): binary arithmetic is performed at the wider of its two operand widths, and a literal operand adopts the other operand's width -- so i8 + 1 is i8 arithmetic, i8 + i16 is i16 arithmetic, and 1 + 2 is i32 arithmetic. signedness is part of the width: there are NO usual arithmetic conversions -- mixing a signed and an unsigned operand in arithmetic or a comparison is an error, cast one side explicitly (a literal adopts the other side, so u64 * 1000 is fine).
wrapping: add, sub, mul and shl wrap at their operand width, and the wrap is DEFINED (signed wrap is two's complement; unsigned wrap is mod 2**width). div and mod have no defined wrap: division by zero traps in the hardware, and INT_MIN / -1 traps too (signed only) -- and when both operands are compile-time constants, these are compile errors instead (fail fast). >> is arithmetic (sar) on signed types and logical (shr) on unsigned; << fills zeros. a CONSTANT shift count outside [0, width) is a compile error; a non-constant count is the programmer's business (the hardware masks it, nsc does not define that).
[h2]comparisons and logic[e]
a < b, a <= b, a > b, a >= b, a == b, a != b produce an i32 (0 or 1); the operands are unified to the wider width first. ptr == ptr and ptr != ptr compare addresses; ordered comparison of ptrs is an error, and ptr == 0 is an error (see null pointer). a && b, a || b and !x produce an i32; their operands are conditions, cast to i64 by jtype3 (so is if/while/switch/?:'s condition). && and || short-circuit, exactly as c89.
[h2]pointers[e]
ptr arithmetic is byte-scaled (there is no pointed-to type to scale by): ptr + int and int + ptr give ptr, ptr - ptr gives i64 (a byte distance), ptr - int gives ptr. *p is an 8-byte load yielding i64 (cast the result down to taste); *p = v stores 8 bytes; &x gives ptr. & and * are inverses on lvalues. every pointer is restrict by construction: jalloc7 gives every value its own stack slot and nothing survives in a register across an instruction.
function pointers are ptrs too -- there is no function-pointer type, and no signature rides along with one. &f, for a function f (defined or only declared), is f's address, typed ptr like any other: it stores in a ptr variable, field or array element, passes and returns as a ptr, and compares with == and != (the same function's address is always the same ptr). the bare name f is not a value -- writing it without the & is an error. &f is not a constant expression, so a global cannot be initialized with it; build a dispatch table at runtime (ptr tbl[3]; tbl[0] = &add; ...).
calling a ptr is written e(args), where e is any expression that types to ptr -- fp(1, 2), tbl[i](x), s.fn(x), p->fn(x), get()(x) -- and never (*fp)(args): *fp already means an 8-byte load yielding i64, so (*fp)(args) is a call through an i64, which is an error. which kind of call a bare name(args) is follows one rule: if name is a function (defined or declared) it is a direct call to it, exactly as it always was; otherwise name must be a variable holding a ptr, and it is an indirect call. so a function name always wins in call position: a local ptr sharing a function's name cannot be called by that bare name (rename it). under &, the reverse holds: a variable always wins, and &name is a function's address only when no variable of that name is in scope. neither rule changes what any name meant before function pointers existed. the callee expression is evaluated first, then the args left to right.
an indirect call gets NO compile-time checking of argument count or argument types, and its result is always an i64. this is deliberate, not an oversight: ptr is untyped, so there is no signature to check against, the same way p->f has nothing but the field name to go on and a dynamic array index is the programmer's business. each arg passes at its own type in the next sysv register, and the result is whatever the callee left in %rax, read as an i64 (cast it down to taste, the same as *p; a void callee leaves garbage there, and a struct-returning callee's result is only its first word). calling through a ptr that does not hold a function's address, with the wrong number of args, or with the wrong arg types is the programmer's bug, the same as dereferencing a bad ptr. two things are still checked, because they belong to the call itself rather than the callee: at most 6 args (the six sysv arg registers), and every arg must be a scalar (an int or a ptr -- no struct or array args through a ptr).
[h2]global variables[e]
a file-scope variable declaration (TYPE NAME; or TYPE NAME = EXPR;, outside any function) follows the same static-by-default rule as functions: global exports it, and its absence means the name is still usable within the file (it typechecks, and codegen references it as NAME(%rip) same as an exported one) but the declaration itself is not a definition.
TYPE NAME = EXPR; -- always a real definition, laid out in .data with EXPR's value (an initlist for a struct/array gives one value per field/element, in order), whether global is present or not -- global only controls whether jemit8 exports it (.globl) for other files to link against; a plain initialized declaration is a genuine c-style file-private static, real storage that just isn't visible outside this translation unit.
global TYPE NAME; (no initializer) -- also a real definition, zeroed, sized exactly for TYPE (struct: nwords*8; array: N*element width; scalar: its own width). emitted as a .comm symbol, not a plain label: this is deliberately a TENTATIVE definition in the classic c sense, meant to coalesce with a stronger (initialized) definition of the same name in another file rather than conflict with it -- one file can declare global i32 counter = 0; while another declares global i32 counter; and both mean the same shared variable, the same way c's tentative-definition rule has always worked. two uninitialized global declarations of the same name across files coalesce with each other the same way (largest size wins).
TYPE NAME; (no global, no initializer) -- a reference only, exactly like a bodyless function prototype. allocates nothing; the real definition (global, initialized or not) is expected in another file linked alongside it. referencing a name that is never defined anywhere is a link-time error, same as calling an undefined extern function.
[h2]functions and calls[e]
everything is static by default; the global keyword is the only way to export a symbol. a decl with no body/init is extern-style: it allocates nothing, may stay unresolved forever, and linking is the user's problem -- unnamed params (_ as placeholder) are legal in decls but not in definitions. calling an undeclared function is an error (jscope2); calling with the wrong number of args is an error (jtype3); more than 6 params or 6 args is an error (sysv has six arg registers and nsc never spills args to the stack). a void function's call cannot be used as a value. all of this is about direct calls, where the callee is a function by name: an indirect call through a ptr (see pointers) has no signature to check, so arity, arg types and void-ness go unchecked there.
return: a bare return is only legal in a void function; a non-void function must END in a return -- the last statement of the body (unwrapping trailing blocks) has to be a ret with a value. falling off the end of a non-void function is a compile error; a side branch that falls off is the programmer's bug and returns whatever %rax held.
main: global i32 main(void) or global i32 main(i32 argc, ptr argv), exactly. void main, non-global main, and any other signature are compile errors. (with -s, main's return becomes the process exit code through nscc's own _start.)
[h2]statements[e]
if/else, while, break, continue, return, blocks, and switch are c89, with two deliberate tightenings: switch does NOT fall through (each case is independent; break jumps past the rest of the chain; duplicate case values are an error), and break/continue outside any loop/switch are errors. ++ and -- exist pre and post on int lvalues, compound assignment exists on int and ptr, and an expression statement must be an assignment, a call, or ++/-- -- anything else is an error.
[h2]sizeof[e]
sizeof(T) and sizeof(expr) are compile-time constants of type i64, folded by jtype3 (sizeof(void) is an error). sizes: i8=1, i16=2, i32=4, i64=8, ptr=8.
[h2]fail-fast, all of it[e]
the compile errors this standard demands, collected: undeclared identifiers and functions; redefinitions; shadowing is fine but a same-scope redeclare is not; break/continue placement; ptr in ordered comparison or ptr==int; ptr in arithmetic with anything but the + - rules; implicit int-to-ptr; i64 literal cast down; constant shift count out of range; constant division by zero; constant INT_MIN / -1; falling off a non-void function; main's signature; void as a variable/expr/param-mate; more than 6 params or args; calling through anything that is not a ptr (including (*fp)(args)); a bare function name used as a value (write &f); a struct or array arg to an indirect call; empty (); // comments; trigraph-free by construction; __ identifiers; the reserved words; unterminated comments and literals; char literals holding more than one char; integer literals that fit no type.