Appendix: compiler architecture¶
This appendix describes the DCC C Compiler toolchain: the compiler dcc, the peephole
optimizer dccpeep, the runtime size reducer dccrtlstrip, and the assembler
and linker - either the host-native
m80c/l80c,
or the Microsoft M80/L80 originals running under ntvcm.
Host-resident tools
dcc, dccpeep, dccrtlstrip,
m80c, and
l80c run on Windows, macOS, and
Linux. They never run on a Z80. They emit Z80 assembly text and CP/M
.COM files that run under CP/M-80, for example via the ntvcm emulator.
m80c/
l80c are clean-room,
LINK-80-object-format-compatible
reimplementations of M80/L80 with unbounded host memory instead of
CP/M's own 64K tables - large nopeep builds can exhaust real L80's own
in-emulator linking workspace well before the target program itself would
not fit; l80c has no such ceiling.
They are the default; pass
dcc-use-emulated-m80=true/dcc-use-emulated-l80=true to dccmake to use
the real M80.COM/L80.COM under ntvcm instead, e.g. to cross-check
output.
The compiler implementation is portable C11 host code built by modern Clang,
GCC, or MSVC. That implementation language is independent of dcc's C89
base language and selected C99/C11 additions for the 16-bit Z80 target.
The toolchain at a glance¶
A single .c file becomes a CP/M .COM executable through a short pipeline.
Each stage has one job and hands a text or object file to the next:
flowchart TB
SRC["C source"] --> DCC["dcc<br/>AST → MIR → Z80 assembly"]
DCC --> PEEP["dccpeep<br/>optional assembly optimization"]
PEEP --> APPSTRIP["dccrtlstrip<br/>remove unreachable app blocks"]
APPSTRIP --> APPASM["m80c<br/>assemble application"]
APPASM --> APPREL["app.REL"]
RTL["DCCRTL.MAC<br/>full runtime"] --> STRIP["dccrtlstrip<br/>keep referenced routines"]
APPSTRIP -. runtime references .-> STRIP
STRIP --> RTLASM["m80c<br/>assemble reduced runtime"]
RTLASM --> RTLREL["RTLMIN.REL"]
APPREL --> LINK["l80c<br/>link application + runtime"]
RTLREL --> LINK
LINK --> COM["CP/M .COM executable"]
| Stage | Tool | Input | Output | Role |
|---|---|---|---|---|
| Compile | dcc |
.c |
.MAC |
Parse typed AST, lower and verify MIR, then select Z80/M80 assembly |
| Optimize | dccpeep |
.MAC |
.MAC |
Local peephole rewriting of the asm |
| Reduce application | dccrtlstrip |
All app .MAC files |
Rewritten app .MAC files |
Remove functions, initialized objects, and module BSS objects unreachable from the final program entry |
| Reduce runtime | dccrtlstrip |
DCCRTL.MAC + app .MAC |
RTLMIN.MAC |
Keep only the runtime routines the app references |
| Assemble | m80c |
.MAC |
.REL |
Object code (relocatable); dccmake uses native m80c by default |
| Link | l80c |
.REL files |
.COM |
Resolve symbols into a CP/M executable; dccmake uses native l80c by default |
The dccpeep stage is optional (dccmake dcc-peep=false skips it).
dccrtlstrip first
computes whole-program reachability across all final application assembly
modules, then uses the reduced application to select runtime blocks.
Compiler shape: front end, AST, and MIR¶
DCC C Compiler is an AST/MIR compiler. The recursive-descent front end builds typed AST nodes for function statements and expressions. Those transient nodes lower into one persistent MIR function before the AST arena is reused. Production function assembly is emitted from generated MIR candidates.
Top-level declarations, global initializers, strings, and data/BSS placement remain table-driven because they are not function-body instructions.
flowchart TB
SRC["C source"] --> FRONT["Front end<br/>preprocess, parse, build typed AST"]
FRONT --> MIR["Persistent MIR<br/>lower statements and record metadata"]
MIR --> ANALYZE["Prepare MIR<br/>repair metadata, check structure,<br/>build CFG, promote objects, transform values"]
ANALYZE --> VERIFY["Verify dominance<br/>independent reachable CFG + PHI-edge checks"]
VERIFY --> ALLOCATE["Plan baseline allocation<br/>solve liveness, assign homes and spills"]
ALLOCATE --> SELECT["Select generated Z80 candidate<br/>try exact structural schedules first"]
SELECT -- exact match --> ASM["Selected .MAC function body"]
SELECT -- exact declines --> GENERAL["Build generated alternatives<br/>rollout, homed, regional, spilled"]
GENERAL --> COST["Choose alternative<br/>candidate-specific allocation + mir-v1"]
COST --> ASM
ALLOCATE -. optional reports .-> SHADOW["Diagnostic shadow models<br/>target constraints + sparse schedule"]
| Classic phase | DCC C Compiler implementation |
|---|---|
| Preprocessing / lexical analysis | Integrated macro engine and next_token lexer in dcc_preproc.c; dcc_pp_expr.c evaluates conditional directives |
| Parsing | Recursive descent builds typed AST nodes against live symbol/type tables |
| Intermediate representation | Persistent per-function MIR with virtual values, typed memory, calls, labels, branches, PHIs, VLA operations, and aggregate copies |
| MIR analysis | Deferred metadata repair, structural checks, CFG construction, object promotion, independent dominance verification, liveness, baseline register homes, and spill-slot planning |
| Instruction selection | Priority exact machine schedules, followed by general rollout, homed, hybrid/regional, and spilled CFG candidates |
| Profitability | After exact scheduling declines, mir-v1 compares generated alternatives using machine instructions/bytes, helper/frame/spill costs, moves, branches, and loop weighting |
| Machine-dependent cleanup | Standalone dccpeep fixpoint optimization over the selected assembly |
Transient AST, persistent MIR¶
AST nodes retain C result type and effective operand type, so integer promotions, pointer scaling, long/float operations, bit-fields, and lvalue widths are decided before lowering. The MIR keeps those decisions after the AST arena is reset.
MIR uses virtual values rather than source-register claims. Its instruction set represents constants, parameters, loads/stores, address/member/index formation, unary/binary operations, call-site-tagged arguments and calls, aggregate copies, labels, branches, jumps, edge-specific PHIs, returns, and VLA stack save/allocate/restore operations.
One production body walk¶
dcc_func.c performs one production metadata/MIR walk for each function.
Declarations are represented by stable placeholders because parser replay may
discover initializers, renamed C99 loop variables, strings, inline temporaries,
or VLA scope events after surrounding statements have already lowered.
dcc_ast_metadata.c and dcc_ast_stmt_meta.c fill and position those events
without generating a second assembly body.
This separation is load-bearing:
- declarations, scopes, VLA exits, labels, diagnostics, and debug events have explicit non-emitting owners;
- statement evaluation occurs once;
- initializer side effects cannot be duplicated by metadata replay;
- every production function body reaches instruction selection as verified MIR.
Verification, promotion, and allocation¶
Structural verification checks operand/object bounds, labels, definitions,
PHI references, call identities, and known argument ABI types. After object
promotion and semantic transformations, dcc_mir_verify.c independently
reconstructs the instruction CFG and computes an immediate-dominator tree in
reverse postorder, using storage linear in the MIR size.
Every ordinary value use on a reachable path must be dominated by its definition: execution cannot reach the use without first passing the definition. PHIs define their results at the logical block entry, even when they appear later in the instruction array; each input must dominate its corresponding incoming edge. Argument records and their values must dominate the matching call. Unreachable paths impose no dominance requirement, but their IDs and references still undergo structural checks. This non-mutating check runs before allocation and candidate emission and cannot be disabled through an environment setting.
Object promotion reuses agreeing reaching definitions or constructs a valid two-predecessor PHI. It distinguishes an undefined entry value from an unreached dataflow state: a definition discovered only on a loop backedge cannot supply the function-entry path. Undefined values and unresolved joins remain memory operations. This does not make reading an uninitialized C local defined behavior; it prevents that read from being represented by an invalid SSA value. Exact schedules account for the corrected NOP positions where unnecessary loop merges no longer produce PHIs.
After verification, backwards liveness treats PHI operands as incoming-edge uses and keeps call arguments live through the matching call-site ID.
Allocation assigns lifetime homes in HL, DE, BC, or callee-saved IY, with deterministic spill slots when pressure or ABI constraints require them. Call-crossing ordinary values may use only IY. Fixed Z80 operand/result registers are boundary constraints, so the emitter inserts moves rather than precoloring a value for its whole lifetime.
The MIR analysis pipeline computes a baseline allocation and retains its liveness matrices for selection. Candidate construction may then derive and measure a different allocation plan: lazy-parameter and regional candidates, for example, save the baseline homes and spills, recompute them for that candidate, and restore the baseline before the next attempt.
Generated-only candidate selection¶
Every candidate writes to its own temporary stream. A declining selector cannot leave partial text in the next candidate. Production first tries:
- exact
scheduled-machine-cfgkernels with complete structural proofs; - the compact
general-rolloutscalar DAG for eligible straight-line functions; homed-scalar-cfg, including hybrid and regional-home variants;- the general
spilled-scalar-cfgemitter.
An accepted exact schedule has priority. When exact scheduling declines, the
selector establishes a complete generated incumbent from the rollout or general
CFG candidates, then mir-v1 may compare that incumbent with homed,
lazy-parameter, hybrid, regional, and spilled variants. The policy compares
only generated MIR candidates.
DCC_MIR_REQUIRE_COMPLETE=1 and DCC_MIR_REQUIRE_EMIT=1 are the strict
semantic and generated-output boundaries.
Proofs are deliberately conservative. Unknown, recursive, cyclic, volatile, aliased, or unsupported shapes decline an optimization, not MIR emission.
Refactor validation¶
For behavior-preserving front-end or metadata refactors, compare raw compiler output and diagnostics before and after, then run the CP/M suite. For MIR selection changes, compare stack and no-stack census snapshots, run every app whose generated hash or metrics changed in peep and nopeep modes, and finish with strict full+extended stack and no-stack runs.
Compiler Features¶
The DCC C Compiler is intentionally small, but the compiler front end still provides a user-facing feature set around diagnostics, C89/C99 compatibility extensions, target-model checks, and size-oriented dead-code elimination in the generated program.
Error messages by category¶
Compiler diagnostics are emitted through dcc_diag_emit.c. Each diagnostic is
assigned a stable DCC-E#### code, reported with the source file and line, and
prints the original source line with a caret when the compiler has an exact
token position. The compile-fail diagnostic tests under tests/diagnostics/
lock these messages and carets against exact baselines.
| Category | Code range | Examples |
|---|---|---|
| Name lookup | DCC-E0201 |
undeclared identifiers |
| Preprocessor and include handling | DCC-E0301-DCC-E0321 |
malformed #include, unknown directives, unmatched #elif/#else/#endif, bad macro arity |
| Constant expressions | DCC-E0401-DCC-E0403 |
non-constant expressions, division by zero, missing expression operands |
| Struct/union/enum/type semantics | DCC-E0501-DCC-E0540 |
field designators, offsetof, bit-fields, duplicate enum constants, missing type names, multiple storage classes |
| Array and pointer constraints | DCC-E0601-DCC-E0605 |
unsupported variable inner dimensions, invalid bounds/object sizes, scalar subscripting |
| Statement control flow | DCC-E0701-DCC-E0706 |
break/continue outside valid contexts, stray case/default, duplicate or undefined labels |
| Functions and declarations | DCC-E0801-DCC-E0806 |
parameter declaration errors, redefinitions, too few or too many function-call arguments |
| Initializers and assignment compatibility | DCC-E0901-DCC-E0921 |
invalid initializers, integer-to-pointer assignment, taking the address of a register-qualified object |
| Unsupported or malformed constructs | DCC-E1001-DCC-E1005 |
unsupported AST forms, malformed syntax, oversized string literals, unsupported sizeof expressions |
| General syntax and top-level parsing | DCC-E1101-DCC-E1107 |
expected tokens such as ;, ), ], =, or an external declaration |
| CP/M/Z80 target-model limits | DCC-E1201-DCC-E1203 |
unsupported double, long long, and 64-bit integer typedef names |
The categories are deliberately broad: the numeric code says where the problem was detected, while the message text names the exact source-level issue.
Dead code detection and elimination¶
The DCC C Compiler does not run a whole-program control-flow analysis pass, and it does not warn about arbitrary unreachable user statements. Instead, dead-code handling is pragmatic and size-focused at the points where the toolchain has reliable local knowledge:
- Dead MIR values and stores. Lowering records side effects separately from virtual results. Liveness, object promotion, and selector-local proofs can omit unused values, dead stores, and rematerializable temporaries.
- Dead stores and labels in assembly.
dccpeepruns to a fixpoint over the emitted.MACtext. Among its cleanups are jump/label threading, jump-to-next removal, redundant load/store removal, and conservative dead IX-frame store elimination when a stack slot is overwritten before it can be read. - Dead static inline bodies. Simple
static inlinefunctions can be buffered and emitted only when a real out-of-line body is needed, avoiding unused helper text in the generated assembly. - Dead runtime blocks.
dccrtlstripperforms conservative mark-and-sweep dead-block elimination overDCCRTL.MAC: it roots symbols referenced by the app, follows runtime-to-runtime references to a fixpoint, and writesRTLMIN.MACcontaining only reachable runtime blocks. - Dead application blocks. Before runtime selection,
dccrtlstripfollows reachability across the final application's marked assembly modules and removes unreferenced functions and eligible objects. This is link-time reachability, not a whole-program C statement analysis.
This split keeps the compiler simple while still attacking the biggest sources of wasted code: unused expression values, local assembly redundancies, unused inline bodies, and unreferenced runtime support.
Inside DCC C Compiler: module architecture¶
The compiler is one binary built from focused modules. Foundational target
types, shared data structures, and broadly used APIs live in dcc.h; narrower
contracts live in subsystem *_internal.h headers. Shared compiler state is
defined once in dcc_state.c; machine-family modules do not add shared data.
graph TB
subgraph FE["Front end"]
DRIVER["dcc.c<br/>driver + CLI"]
PP["dcc_preproc.c / dcc_pp_expr.c<br/>preprocessor + lexer"]
PARSER["dcc_func.c / dcc_stmt.c<br/>declarations + statements"]
SEM["dcc_types.c / dcc_symbols.c<br/>types + symbols"]
end
subgraph AST["Transient typed AST"]
BUILD["dcc_ast_build.c"]
NODES["dcc_ast.c / dcc_ast.h"]
META["dcc_ast_metadata.c<br/>dcc_ast_stmt_meta.c"]
end
subgraph MIR["Persistent function MIR"]
CORE["dcc_mir.c<br/>lowering + repair + promotion"]
VERIFY["dcc_mir_verify.c<br/>independent dominance verification"]
ALLOCATE["dcc_mir.c<br/>liveness + baseline allocation"]
SHADOW["dcc_mir_target.c / dcc_mir_schedule.c<br/>diagnostic shadow models"]
end
subgraph BACK["Generated back ends"]
SELECT["dcc_mir_select.c<br/>priority, rollout + mir-v1"]
COMMON["dcc_mir_emit_common.c<br/>scalar DAG + shared emission"]
HOMED["dcc_mir_homed_cfg.c<br/>homed / hybrid / regional"]
SPILLED["dcc_mir_spilled_cfg.c<br/>general spilled CFG"]
MACHINE["dcc_mir_machine_*.c<br/>exact structural schedules"]
end
subgraph OUT["Top level and data"]
INIT["dcc_global_init.c<br/>file-scope initializers"]
DATA["dcc_data.c<br/>strings, globals, BSS"]
ASM["selected function streams<br/>+ top-level assembly"]
end
DRIVER --> PP --> PARSER --> BUILD --> CORE
SEM --> BUILD
NODES -.storage and typed nodes.-> BUILD
META -.declarations / scopes / VLA events.-> CORE
CORE --> VERIFY --> ALLOCATE --> SELECT
ALLOCATE -. diagnostic reports .-> SHADOW
SELECT --> COMMON
SELECT --> HOMED
SELECT --> SPILLED
SELECT --> MACHINE
HOMED -. shared operations .-> COMMON
SPILLED -. shared operations .-> COMMON
COMMON --> ASM
HOMED --> ASM
SPILLED --> ASM
MACHINE --> ASM
INIT --> DATA --> ASM
| Group | Modules | Responsibility |
|---|---|---|
| Shared | dcc.h, dcc_state.c, subsystem *_internal.h files |
Target model, shared contracts, and lifecycle-owned compiler state |
| Front end | dcc.c, dcc_preproc.c, dcc_pp_expr.c, dcc_func.c, dcc_stmt.c, dcc_diag_emit.c, dcc_global_scan.c |
Driver, preprocessing/lexing, declarations/statements, conservative global-use prepass, frame scan, and diagnostics |
| Types / symbols | dcc_types.c, dcc_symbols.c, dcc_constexpr.c, dcc_fold.c, dcc_asmname.c |
Type system, symbol tables, constant evaluation/folding, and M80-safe assembly-name mapping |
| Typed AST / metadata | dcc_ast.c, dcc_ast_build.c, dcc_ast_gen*.c, dcc_ast_metadata.c, dcc_ast_stmt_meta.c, dcc_licm.c |
Transient typed trees, semantic classifiers, and non-emitting LICM/CSE planning and declaration/scope replay |
| Compatibility helpers | dcc_expr.c, dcc_ops.c, dcc_cmp.c, dcc_assign.c, dcc_decl.c, dcc_stmt_fast.c, dcc_array_narrow.c |
Shared initializer/type behavior and conservative source proofs; not a production body emitter |
| MIR core | dcc_mir.c, dcc_mir_stream.c |
Persistent IR, metadata repair, CFG/verifier, liveness, baseline allocation, and isolated candidate streams |
| MIR dominance verification | dcc_mir_verify.c |
Independent reachable CFG, immediate dominators, ordinary-value and PHI-edge checks, and call-argument dominance; no IR rewriting or allocation |
| MIR emission / selection | dcc_mir_select.c, dcc_mir_emit_common.c, dcc_mir_homed_cfg.c, dcc_mir_spilled_cfg.c |
Exact-schedule priority, scalar DAG rollout, transactional generated candidates, candidate-specific homes/spills, shared emission, and mir-v1 selection |
| Diagnostic shadow models | dcc_mir_target.c, dcc_mir_schedule.c |
Optional Z80 constraint and sparse-schedule reports; these modules do not select production candidates or emit Z80 |
| Machine schedules | dcc_mir_machine_emit.c (coordinator), dcc_mir_machine_*.c (families) |
Exact structural matchers and specialized Z80 streams |
| Top level / output | dcc_func.c, dcc_global_init.c, dcc_data.c |
Function/frame parsing, one production metadata/MIR body walk, global initializer recording, deferred static-body placement, and data-section emission |
Exact machine-schedule families¶
dcc_mir_machine_emit.c owns common machine helpers and the order-sensitive
dispatch. Cohesive schedules live in separately compiled families:
| Family module | Responsibility |
|---|---|
dcc_mir_machine_attention.c |
Matrix, attention, and fixed-point kernels |
dcc_mir_machine_byte_scans.c |
Byte/row scans, fills, copies, hashes, records, and file-line kernels |
dcc_mir_machine_constant_folding.c |
Constant/result flows, result switches, and indexed-member schedules |
dcc_mir_machine_containers.c |
Array, container, stack, comparison, and reduction schedules |
dcc_mir_machine_float_recursion.c |
Floating-point, recursive, tree, and byte-status kernels |
dcc_mir_machine_numeric.c |
Integer, long, fixed-point, and math kernels |
dcc_mir_machine_float_reports.c |
Float reports and checks |
dcc_mir_machine_scanners.c |
Scan, parse, and traversal loops |
dcc_mir_machine_aggregate_checks.c |
Aggregate, array, and struct checks |
dcc_mir_machine_structural_checks.c |
Literal, bitset, sieve, string, structure, and bitfield validations |
dcc_mir_machine_runtime_runners.c |
Runtime, file, and system orchestration |
dcc_mir_machine_interpreter_runners.c |
Interpreter and parser runners |
dcc_mir_machine_call_runners.c |
Call/control orchestration |
dcc_mir_machine_validation_runners.c |
Scope, wide-value, and validation runners |
dcc_mir_machine_wide_records.c |
Wide arithmetic, aggregate updates, and record-oriented schedules |
dcc_mir_machine_endgame.c |
Large final exact schedule families |
Machine families follow a zero-shared-state rule:
- plans, candidates, and mutable matching state are automatic and attempt-local;
- each family exports one dispatcher and no data;
- module-local constant tables use internal linkage;
dcc_mir_machine_internal.hcontains function contracts, not shared plan storage.
Run scripts/audit-c-module-exports.py after changing a module boundary. The
audit must report only the allowed dispatcher and no exported data.
Inside dccpeep: a fixpoint peephole optimizer¶
dccpeep is DCC C Compiler's machine-dependent optimizer. It reads the emitted .MAC
as an array of text lines and applies dozens of small peephole rewrites —
each one matches a short local instruction pattern and replaces it with a
cheaper equivalent (for example folding a redundant store/reload, threading a
jump-to-jump, turning an ld/cp against zero into or a, or replacing an
absolute jp in range with a relative jr).
The optimizer is split by responsibility. peep_lines.c owns storage,
mutation, physical-line I/O, and opaque barriers for ; dcc user asm regions;
passes never rewrite or delete those user-authored lines. peep_parse.c
contains stateless assembly parsers, peep_effects.c caches structured line,
opcode, register, and flag metadata, and peep_analyze.c contains shared
conservative register analysis. peep_control_flow.c owns versioned
label/function indexes and bounded reachability queries. The main dccpeep.c
file owns the descriptor-driven, order-sensitive fixed-point catalogue and
remaining general passes. PeepContext groups program ownership, options,
statistics, mutation versions, and cached indexes; compatibility globals
preserve established pass signatures. PeepEditTransaction provides opt-in
atomic commit/rollback for coupled rewrites.
The high-volume local dispatcher, board/game idioms, loop registerization, and
compiler-tagged temporary handling live in peep_pass_once.c,
peep_pass_minmax.c, peep_pass_loops.c, and
peep_pass_inline_temp.c. Label and branch rewrites live in
peep_pass_control_flow.c. Post-convergence shared-helper rewrites and
terminal cleanup live in peep_pass_stubs.c and peep_pass_final.c
respectively, behind dccpeep_internal.h.
dccpeep runs its rewrite catalogue to a fixpoint so one rewrite can expose the pattern another rewrite needs:
flowchart TB
READ["read .MAC into line array"] --> SWEEP["run all peephole passes once"]
SWEEP --> CHK{"any pass<br/>changed a line?"}
CHK -->|yes, and under 30 passes| SWEEP
CHK -->|no change, or 30 passes| POST["post-convergence passes"]
POST --> FRAME["IX-frame elimination<br/>+ optional shared stubs (-Os)"]
FRAME --> FINAL["final tidy:<br/>dead-load removal, jp -> jr"]
FINAL --> WRITE["write optimized .MAC"]
Key design points:
- Driven to convergence. The main loop re-runs every pass until a full sweep makes no change (capped at 30 iterations), because passes feed each other — e.g. a store/reload elimination can create a dead label that a later pass then removes.
- Order-sensitive post-passes. A few transforms must run after convergence: the signed-compare constant-bias fold, IX-frame elimination, and the shared-stub passes all depend on the instruction stream having settled, because earlier structural passes recognise loops by their canonical (un-folded) shape.
- Two optimization goals.
-Ot(default, "time") inlines helper sequences for speed;-Os("size") factors recurring sequences into sharedcallstubs to shrink the binary. Each helper family counts complete matches first and fires only at its whole-program linked-stub break-even. The mode is chosen on the command line and changes which post-convergence passes fire. - Text safety is explicit. Physical input lines are read without a fixed length limit. Marked user-assembly regions become opaque label barriers while optimization runs, then are restored byte-for-byte; address relaxation also refuses to estimate across them.
- Direct observability.
scripts/run-dccpeep-tests.ps1checks focused default,-Os, and undocumented-opcode fixtures plus idempotence. Optional-fstatsoutput reports convergence iterations, per-pass calls and changes, and inserted/deleted line totals to stderr without changing normal output.
Why a separate program instead of an in-compiler pass
Keeping the peephole optimizer as a standalone text-to-text filter keeps
the compiler simple and allows inspection, diffing, and skipping
optimization (nopeep). The assembly is the contract between the two tools.
The runtime: a block-structured library sized for stripping¶
The runtime DCCRTL.MAC is a single assembly source, but its
architecture is what makes the toolchain's "pay only for what you use"
property possible. Rather than one monolithic blob, the runtime is written as
hundreds of parsed blocks around public entry points and shared preludes. A
program never links the whole library — dccrtlstrip keeps only the blocks the
application actually references (the mark-and-sweep details are in the companion
appendix Runtime optimization). The architectural
consequence is that routine dependencies can be inspected and their linked
cost measured for a particular application.
flowchart TB
subgraph RT["DCCRTL.MAC (block-structured runtime)"]
BASE["always-present baseline<br/>(start, argv/console, heap, exit)"]
IO["stdio blocks<br/>printf, file I/O core"]
MEM["memory blocks<br/>malloc/free/realloc"]
LONG["32-bit long blocks"]
FLT["float core + math.h"]
STR["string / ctype blocks"]
end
BASE --> APP(["linked .COM<br/>baseline + referenced blocks only"])
IO -. if referenced .-> APP
MEM -. if referenced .-> APP
LONG -. if referenced .-> APP
FLT -. if referenced .-> APP
STR -. if referenced .-> APP
Two numbers describe every routine¶
Because the runtime is block-structured, each public routine has two costs that
the build-time size hook (docs/docs/hooks/runtime_sizes.py) measures directly:
- self — the source lines in the routine's own block.
- marginal — self plus every additional block it transitively links beyond the always-present baseline. This estimates source volume, not assembled bytes or execution cycles.
The gap between the two is the whole story of the runtime's size architecture: a
small self with a large marginal means the routine sits on top of a big
shared substrate (the file-I/O core, or the float arithmetic core).
The always-present baseline¶
The forced start root retains entry, heap/BSS setup, heap-state words, exit,
and their dependencies. Command-tail construction and console output have
separate public blocks; the application main shim can retain additional
support. Use the generated table for the current baseline rather than a fixed
block count. putchar and puts are small wrappers, not zero-cost operations.
What the feature groups cost¶
The runtime's size is dominated by a few shared cores. Routines that sit on a core are cheap individually but expensive to introduce, because the first one links the whole core:
| Feature group | Shared core it tends to link |
|---|---|
Console output (putchar, puts, integer printf) |
baseline console writer or self-contained formatter |
File-stream + low-level I/O (fopen, fread, fputs, fprintf) |
FCB/DMA file core |
scanf / sscanf / fscanf |
shared scan core |
Memory (malloc/free/realloc/calloc) |
heap helpers and size arithmetic |
32-bit long arithmetic |
long multiply/divide/modulo helpers |
float operators |
float normalise/round core |
math.h (sinf, expf, powf, …) |
float core plus conversions and chained math helpers |
string.h / ctype.h |
usually self-contained routines |
The exact per-routine self/marginal numbers — and the transitive
dependencies behind each one — are tabulated on the auto-generated
Runtime function sizes page, with the optimisation
takeaways in Runtime optimization. The takeaways for
sizing a program are:
- Console-only output and string/ctype routines are cheap — they link little or nothing beyond the baseline.
- The first file,
scanf, orfloatoperation is the expensive one; it links a shared core. Additional routines in the same family are then nearly free. math.his the single biggest lever — theexp/log/powand hyperbolic group are expensive because each chains other math routines on top of the float core.
Those numbers are recomputed from DCCRTL.MAC on every docs build, so editing
the runtime and rebuilding the docs is all that is needed to refresh them.
Architecture summary¶
- DCC C Compiler is an AST/MIR C89-base compiler with selected C99/C11 additions. Typed AST nodes lower into persistent verified function MIR.
- MIR owns CFG, PHIs, virtual values, object promotion, liveness, register homes, spills, target constraints, and generated candidate selection.
- Production function assembly comes only from MIR.
- Structurally proven exact machine schedules have priority. When they decline,
general rollout and homed/spilled CFG candidates are selected from generated
streams, with
mir-v1arbitrating eligible alternatives. - Machine-dependent optimization is split out into
dccpeep, a fixpoint peephole optimizer over the assembly text, with separate time (-Ot) and size (-Os) strategies. - The runtime
DCCRTL.MACis block-structured into hundreds of public blocks, so every routine has a measurableself/marginalsize cost anddccrtlstripcan link only the blocks a program references. The exact current totals are generated on the runtime size page. - The back half of the pipeline uses native
m80c/l80c(or MicrosoftM80/L80underntvcm) for assembly and linking - a shared LINK-80-compatible.RELobject format that DCC C Compiler consumes as a fixed target rather than needing to invent its own.