CI Pipeline¶
Mental model for the CI workflow under .github/workflows/ — what runs in what order, what fails fast, and what to regenerate locally before pushing.
Workflow shape¶
┌────────────────────────────────────────────────────────────────────┐
│ Quick gates (parallel, ~30s each) │
│ │
│ ┌──────────────┐ ┌─────────────────────┐ ┌─────────────────┐ │
│ │ format-check │ │ verify-docs-current │ │ codegen-check │ │
│ └──────────────┘ └─────────────────────┘ └─────────────────┘ │
└────────────────────────────────────────────────────────────────────┘
│
needs: [format-check, verify-docs-current]
│
▼
┌────────────────────────────────────────────────────────────────────┐
│ OS build matrix (parallel, ~5–10 min each) │
│ │
│ ┌──────────┐ ┌───────────┐ ┌───────┐ ┌─────────────┐ ┌──────────┐ │
│ │linux-gcc │ │linux-clang│ │ macos │ │windows-msvc │ │win-clang │ │
│ └──────────┘ └───────────┘ └───────┘ └─────────────┘ └──────────┘ │
│ ┌────────────────────┐ ┌────────────────────────┐ │
│ │sanitizers (address)│ │sanitizers (undefined) │ │
│ └────────────────────┘ └────────────────────────┘ │
└────────────────────────────────────────────────────────────────────┘
codegen-check runs in a separate workflow (codegen.yml) — it's not in the dependency chain because it's already fast (~30s) and runs in parallel.
Fast-fail wiring¶
Two mechanisms keep the pipeline cheap on bad pushes:
needs:dependency. Every OS build job inci.ymldeclaresneeds: [format-check, verify-docs-current]. If either quick gate fails, the multi-OS matrix is skipped — saving ~50 compute minutes (10 min × 5 jobs).concurrency.cancel-in-progress. A new push to the same branch / PR cancels the in-flight workflow run for that ref. No more "two builds racing on stale code".push:is filtered tomain. With an unfilteredpush:alongsidepull_request:, every push to a branch with an open PR ran the whole matrix twice, and threeclang-formatjobs becausebuild-matrix.ymlduplicatedci.yml's. A branch now gets exactly one run, frompull_request.
Both are at the workflow top:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
linux-gcc:
needs: [format-check, verify-docs-current]
...
What runs in each gate¶
format-check¶
Just ./scripts/check-format.sh — runs clang-format --dry-run -Werror on every C/C++/header file. Ten seconds. Failures show the exact lines that differ from the configured style; fix is clang-format -i <file>.
verify-docs-current¶
Runs eight checks in sequence (each takes <5s):
| Step | What it does | If it fails |
|---|---|---|
gen_indicator_docs.py |
Indicator reference matches registry.def |
Run python3 scripts/gen_indicator_docs.py |
gen_llms_txt.py --check |
docs/llms.txt + llms-full.txt match docs/ |
Run python3 scripts/gen_llms_txt.py |
check_dts_exports.py |
node/index.d.ts matches NAPI exports |
Edit .d.ts to add/remove the listed names |
check_binding_parity.py |
pybind11/NAPI/Codon/QuickJS coverage matches IDL | See parity-gate.md |
check_error_codes.py |
Every error code has a doc page; pages aren't stale | Add the doc page or remove the unused code |
check_test_gating.py |
Every tests/*.cpp is registered in CMake |
Register the target — see test-gating.md |
check_suite_discovery.py |
Every Python/Node test file is reachable by the suite runners, and nobody re-listed files by hand | Rename the file to test_*, add cases, or drop the hand-written step |
check_quickjs_registration.py |
Every __lrvx_* global the JS layer calls has an addGlobalFunc registration |
Register the global in js_bindings.cpp |
gen_api_index.py --check |
docs/reference/python/_api_index.md matches .pyi |
Run python3 scripts/gen_api_index.py |
check_doc_snippets.py |
Doc snippets follow --8<-- include pattern |
Refactor inline snippets into includes |
sync_mcp_data.py --check |
mcp/lrvx_mcp/data/ matches source |
Run python3 scripts/sync_mcp_data.py |
check_sanitizer_scan.py |
Every job that builds with -fsanitize= runs its tests through the transcript scan |
Wrap the step in scripts/run-with-sanitizer-scan.sh |
| lrvx-mcp pytest | MCP server unit tests, with the compiled binding required | Fix the broken test, or build _lrvx if the step reports a missing dependency |
linux-gcc — the binding suites¶
The one job that builds every binding, so it is where the binding-level checks live. All of them run the whole suite, not a list:
| Step | What it does | If it fails |
|---|---|---|
pytest python/tests -q |
The entire Python suite, under PYTHONPATH=build/python |
Fix the test. Reproduce locally with scripts/ci-local.sh, which uses the same PYTHONPATH — a relative path, which is why a subprocess spawned with a different cwd must absolutise it |
for f in test/test_*.js |
Every Node test file | Fix the test |
check_binding_smoke.py |
Calls ~760 zero-argument methods across both bindings and fails on binding-level errors (unregistered return type, signature the addon cannot accept) | The wrapper is broken, not the input — see parity-gate.md |
Both suites used to be invoked as one named step per file — roughly 35 of them.
A test file nobody remembered to add simply never ran: 34 of 69 Python files and
9 of 21 Node files, including every regression gate written for the
binding-surface audit. check_suite_discovery.py now fails CI if a test file
would not be picked up, or if anyone re-lists files by hand.
codegen-check (separate workflow)¶
Re-runs the IDL→header/codon/markdown emitters and verifies the output matches what's committed. If you edited lrvx_capi_spec.hpp and forgot to re-run regenerate.sh, this fails.
OS build matrix¶
Builds the full project (engine + C ABI + tests + benchmarks + Python + Node + Codon + QuickJS), runs ctest, runs all the integration tests, runs cross-binding parity tests (Python ↔ Node, same C++ math), and exercises example programs.
Green does not mean checked¶
Two steps in this pipeline used to report success while checking less than their names implied. Both are fixed. The shapes are worth knowing because they recur.
A sanitizer report inside a passing test¶
ctest --output-on-failure prints the output of the tests ctest decided had
failed. A test whose assertions all pass, but which printed a sanitizer report
on the way, is recorded as Passed and its output never reaches the log. The
undefined-behavior sanitizer does that by default: it prints
runtime error: ..., carries on, and the process exits 0.
That is the case sanitizers are run for. Twice during the 2026-09 audit a real defect in production code turned up exactly that way, under fully green assertions: a data race on the logger's file descriptor, and an out-of-bounds vector read in the execution simulator. Both were found by hand, by re-running with full output and grepping a few thousand lines of transcript.
The run reads its own transcript now. Every sanitizer job calls the suite
through scripts/run-with-sanitizer-scan.sh.
It runs the command, keeps its exit code, and also scans ctest's per-test log —
build/Testing/Temporary/LastTest.log, which holds every test's output
whatever its verdict — for Address, Leak, Thread, Memory and
UndefinedBehavior report banners. A report fails the step even when the command
returned 0, and the failure names the test and the verdict ctest gave it.
Locally, same call:
scripts/check_sanitizer_scan.py keeps the wiring in place: a job that
compiles with -fsanitize= and does not wrap its test step fails the docs
gate. The matcher has a --self-test that feeds it every report shape it
claims to catch, plus text it must not match — an unwatched detector is not a
detector.
A parity gate that skipped a quarter of what it named¶
check_binding_parity.py is the cross-binding parity gate and the project has
four bindings: pybind11, NAPI, Codon, QuickJS. The per-group loop called the
first three. The quickjs key existed in the manifest and was read by no line
of code, and the closing message said "all bindings in parity". A group with no
QuickJS implementation at all passed, which is how a composite-book function
shipped missing from QuickJS with this green.
QuickJS is read now, against the addGlobalFunc registration table in
src/quickjs/js_bindings.cpp — the only list of what a strategy can actually
call. A quickjs: required entry names the globals rather than deriving them,
because the __ plus C-API-name convention has real exceptions
(__lrvx_vprofile_create wraps lrvx_volume_profile_create).
Most groups still carry no quickjs entry, so the gate cannot demand one yet.
It counts them and prints the number instead of implying they were checked, and
the closing line now names what it checked per binding. --require-quickjs
turns undeclared groups into failures; turn it on in CI once the manifest is
filled in.
A test suite that quietly shrank¶
pytest mcp/tests/ reports 229 passed, 20 skipped without the compiled
lrvx binding and 248 passed, 1 skipped with it. Both exit 0, and the 19
cases that drop out are the ones that touch real code rather than fixtures.
Nothing separated the two runs but the numbers.
mcp/tests/conftest.py probes the dependencies
that shrink the suite — the binding and the mcp SDK — and closes the run with
a banner naming each missing one and how many cases it cost. Where the run is
meant to be complete, a missing dependency fails it instead of reducing it:
that is the default under CI, and LRVX_MCP_REQUIRE_DEPS=1 / =0 forces it
either way. A contributor who has not built the C++ side still gets the
pure-Python half, and the count of what did not run.
The docs sync chain (eight scripts, in order)¶
The "verify-docs" gates each check that a generated artifact matches what's committed. Several of those artifacts depend on each other — regenerating one re-derives the next:
1. tools/codegen/scripts/regenerate.sh → lrvx_capi.h, golden/, .api/, mcp data
2. cmake --build build → liblrvx_capi + Python module
3. scripts/gen_pyi_stubs.py → .pyi from running pybind11
4. scripts/gen_api_index.py → docs/reference/python/_api_index.md from .pyi
5. scripts/gen_llms_txt.py → docs/llms.txt + llms-full.txt (embeds api_index)
6. scripts/sync_mcp_data.py → mcp/lrvx_mcp/data/ (handled by regenerate.sh, but run manually if you skipped step 1)
7. scripts/gen_indicator_docs.py → docs/reference/codon/indicators.md (only if you touched registry.def)
8. python3 scripts/check_binding_parity.py → manifests vs bindings
Skipping any one of those usually means a CI failure that takes ~30s to detect, but the order matters — gen_llms_txt.py reads _api_index.md, so regenerating _api_index.md after generating llms-full.txt leaves them out of sync.
If you've touched the IDL spec or any pybind11/NAPI code, run them all in order before committing. If you've only touched a C++ engine internal that doesn't change the public surface, you don't need to.
What fails first (debugging guide)¶
When CI is red, look at the topmost failed step:
- format-check fails → run
clang-format -ion the listed files - codegen-check fails → run
bash tools/codegen/scripts/regenerate.sh - verify-docs-current fails → look at the specific step name; run the script it names
- OS build fails → reproduce locally with the matching toolchain (most issues are platform-specific compile errors visible in the log)
- OS build but only on sanitizers → the bug exists, sanitizers caught what regular tests missed; address it, don't disable the sanitizer
The fast-fail wiring means if quick gates fail, build matrix is SKIPPED (gray, not red) — which is the design. If you see all build jobs gray and only one quick gate red, fix that gate; everything else will run on the next push.
Adding a new gate¶
To add a new check that should block builds:
- Add a step under
verify-docs-current(if it's a docs/sync check) orformat-check(if it's a static check on source). - Make sure the script exits non-zero on failure with a clear
::error::annotation. - Verify the build jobs already
needs: [format-check, verify-docs-current]— they do — so failures will short-circuit the matrix automatically.
If a check needs to run on every OS (e.g. exercising the platform-specific .dylib / .dll), put it in each OS build job instead. The fast-fail logic still applies — it won't even start if quick gates failed.