Gates
Gates are automated quality checks that must pass before a WU can be completed. They replace manual code review with consistent, automated enforcement.
Why Gates?
Section titled “Why Gates?”Traditional review:
- Human bottleneck (waiting for reviewers)
- Inconsistent (different reviewers, different standards)
- Slow feedback (review happens after code is written)
Gates:
- Instant (run automatically)
- Consistent (same checks every time)
- Fast feedback (run locally before pushing)
Config-Driven Gates
Section titled “Config-Driven Gates”Define your gate commands in workspace.yaml under software_delivery.gates:
Optional Migration Verification
Section titled “Optional Migration Verification”Projects with manual database deploy steps can add an explicit migration-state verifier:
When migration_verify is configured, pnpm gates and pnpm wu:prep run it only when the
working diff touches schema or migration paths such as db/schema/**, prisma/schema.prisma,
supabase/schema.sql, or migration directories.
Use this for commands that check whether the target database is up to date. Do not use it to run migrations automatically.
This approach works with any language and toolchain.
Gate Execution Concurrency
Section titled “Gate Execution Concurrency”Gate runs from worktrees in the same repository share a FIFO execution lock. The default capacity
is 2, allowing two isolated worktree runs while keeping host use bounded. Repositories whose gate
commands write shared outputs can opt down to strict serialization:
concurrency must be an integer from 1 through 16. When all slots are occupied, later runs
remain queued in request order. Queue diagnostics report the configured capacity, active holder
identities, and the current WU’s queue position.
Fresh lifecycle CLI builds do not occupy this gate semaphore. They use a distinct ownership-fenced lock scoped to the canonical output checkout: builds for different worktrees can overlap, while a second writer to the same output waits and reuses the first successful result.
Build-output-root serialization
Section titled “Build-output-root serialization”The repository semaphore limits host load, but it is not the authority for shared build output.
Every gate run that may write build artifacts also acquires a capacity-one lock under
~/.lumenflow/locks/, keyed by the canonical build-output root that run writes. Two runs whose
roots are disjoint proceed concurrently; two runs that share a root queue instead of interleaving.
The lock remains held for the complete build-producing run, not one gate at a time, preventing
another run from replacing artifacts between a build and a later test.
The key resolves where output really lands rather than where the checkout sits. Every output root of the run checkout is resolved with the platform’s native realpath, so symlinks, junctions, and case-variant spellings collapse to one identity:
- every resolved output inside the run checkout → the key is that checkout;
- outputs resolving into one other checkout (for example a worktree whose build directory links into another checkout) → the key is that other checkout, and the two runs serialize;
- no resolvable answer → the run falls back to the machine-wide lock at
~/.lumenflow/locks/build-output.lockand serializes exactly as it did before root keying. This covers a run outside any checkout, an output that cannot be resolved, outputs scattered across several other checkouts, a declaration resolving to several distinct roots, aworkspace.yamlthat is present but cannot be loaded, and a scan that found no output directory at all — an empty enumeration proves nothing about where the run writes. Each fallback logs the reason it fired.
Detection recognizes dist, build, out, .turbo, and .next inside the run checkout, and
resolves any directory that links out of the checkout whatever it is named. Any other ecosystem’s
output tree participates by declaring it. On Windows, native realpath collapses symlinks and
junctions but does not collapse subst drives or mapped network drives, so two spellings reached
through those mechanisms key separately and must be declared to serialize.
A workspace whose gate commands write a tree that other checkouts also write declares that tree. Declaring it is the supported way to participate in serialization across checkouts or repositories: every run of the workspace keys on the declaration, so all checkouts whose declarations resolve to an identical root list serialize with one another. One lock directory names exactly one serialization domain, so a declaration resolving to several distinct roots falls back to the machine-wide lock instead of being keyed.
Entries may be absolute or relative to the checkout root. The default is an empty list, which keys each run on its own resolved output root. Parallelism on one machine is therefore bounded by declared shared roots and by available compute through the repository semaphore, never by a machine-wide count.
A project’s own build tooling may hold a separate per-checkout build lock; that lock keeps its role and is neither replaced nor weakened by this one.
Only the built-in read-only allowlist bypasses this lock: format:check, spec:linter,
backlog-sync, claim-validation, supabase-docs:linter, and co-change. Their existing parallel
pre-pass is unchanged. Every other built-in gate is treated as build-producing, and a
consumer-defined gate participates automatically when you register it through the normal gate
configuration—no separate lock flag or shell convention is required.
Automated Test Diff Evidence
Section titled “Automated Test Diff Evidence”wu:prep can enforce an extra proof step: if a WU changes code, it must also touch at least one
automated test file in the same diff. This policy is configured under
software_delivery.gates.tdd_diff_evidence.
Defaults come from software_delivery.methodology.testing:
tddsetssoftware_delivery.gates.tdd_diff_evidence.mode: blocktest-aftersetssoftware_delivery.gates.tdd_diff_evidence.mode: offnonesetssoftware_delivery.gates.tdd_diff_evidence.mode: offapplies_to_typesdefaults tofeatureandbugexempt_pathsdefaults to[]test_file_patternsdefaults to the built-in polyglot conventions listed belowcode_file_extensionsdefaults to TS/JS extensions (.ts,.tsx,.js,.jsx,.mjs,.cjs,.mts,.cts)
mode supports block, warn, and off. wu:prep only blocks when the mode is block; teams
that want the policy disabled can set warn or off and still document their intent in config.
What counts as a test file
Section titled “What counts as a test file”One matcher decides test-ness for both the tdd_diff_evidence gate and the commit-order
(RED-first) gate. It recognises these path conventions by default, with no configuration:
| Ecosystem | Recognised paths |
|---|---|
| TypeScript/JavaScript | **/*.test.{ts,tsx,js,jsx,mjs}, **/*.spec.{ts,tsx,js,jsx,mjs}, **/__tests__/**, **/*.test-utils.*, **/*.mock.* |
| Go | **/*_test.go |
| Python | **/test_*.py, **/*_test.py |
| Rust | **/tests/**/*.rs (Cargo integration tests) |
Production files whose path merely contains the substring test (for example
src/lib/testimonials.ts or src/contest/scoring.go) are not test files.
Explicitly unsupported: Rust inline #[cfg(test)] modules. They live inside the production
source file and carry no distinguishing path, so path-based matching cannot see them. Put Rust tests
in a tests/ directory, or declare the source file under the WU’s tests.unit list so the
tdd_diff_evidence gate can count it.
Configuring for other toolchains
Section titled “Configuring for other toolchains”test_file_patterns and code_file_extensions make the gate language-agnostic. test_file_patterns
is additive: the globs you supply extend the built-in conventions and can never disable them, so
configuring one language never silently stops another language’s tests from counting. There is no
supported way to reduce recognition below the defaults. code_file_extensions still replaces the
default extension list.
When the commit-order gate rejects a commit with test files in same commit: 0, its output lists
every recognised convention so an unrecognised ecosystem is visible rather than silent.
C# (xUnit / NUnit / MSTest):
Python (pytest / unittest) — test paths are recognised by default; only the production extension
needs configuring, plus any extra convention such as a shared tests/ tree:
Go — *_test.go is recognised by default:
Use this gate when you want changed-test evidence for specific WU types or runtime paths. It is not a requirement to force every team into test-first development.
For one-off exceptions, document the reason in the WU notes with:
Native Delivery Review
Section titled “Native Delivery Review”The Software Delivery pack also exposes a native delivery_review gate for completion review. This
capability is public and vendor-agnostic:
- It lives under
software_delivery.gates.delivery_review enabled: trueregisters it for every agent/client runtimeauto_run: truemakeswu:preprun it for applicable WU types- It runs through native gate execution and
wu:prep, not through a vendor-specific skill path - It does not depend on lumenflow-cloud or any hosted control plane
pnpm gates runs delivery_review whenever the global gate is enabled. pnpm wu:prep auto-runs
it when both enabled: true and auto_run: true are set, unless the current WU type matches
skip_types. The same public contract is catalogued in the
Gates Reference.
For source-code delivery changes, delivery_review requires automated test evidence or meaningful
manual verification evidence. A non-empty tests.manual entry is not enough by itself:
placeholders and negative values such as todo, n/a, screenshot: n/a, and not run are
treated as missing evidence and fail the gate. Use concrete manual entries that name the surface,
action, observed result, and artifact path when visual or manual QA is the right evidence.
Set block_partial: true when the repo wants delivery-review uncertainty to block rather than
warn. Set verifier_command when native evidence-shape checks should be followed by project-owned
truth checks. The command runs from the repository root, receives
LUMENFLOW_DELIVERY_REVIEW_WU_ID=<WU-ID>, and blocks the gate when it exits non-zero.
verifier_command_mode (WU-3982) controls when that command runs: always (the default) runs it
every time one is configured, matching every prior release byte-for-byte. on-insufficient-evidence
skips the run when native evidence — automated test evidence, meaningful manual evidence, or a
documented exemption — already reports sufficient evidence for the change set, and the gate log
states the verifier was skipped for that reason. Use it to keep an expensive consumer-owned
verifier (such as a Playwright visual harness) from running on a change a component test already
proves.
Client-specific config can still exist for adapter UX, hooks, or prompt surfacing, but it is not the enforcement switch:
For one release, legacy client-scoped features.delivery_review.auto_run: true is still honored
when the global gate is enabled and global auto_run is omitted. LumenFlow emits a migration
warning pointing to software_delivery.gates.delivery_review.auto_run. Client-scoped
features.delivery_review.enabled: false does not disable the core gate; disable or skip the gate
through the normal global gate config or auditable gate-skip mechanism.
Delivery Review Output Contract
Section titled “Delivery Review Output Contract”delivery_review produces a stable JSON artifact at
.lumenflow/artifacts/delivery-review/<WU-ID>.json. Hosts and products can consume the result
without assuming a specific vendor runtime.
Verdict behavior:
PASSmeans the review found sufficient delivery evidencePARTIALmeans the review completed with uncertainty or lower-severity findings; it warns by default and blocks whensoftware_delivery.gates.delivery_review.block_partialistrueFAILmeans the review found blocking gaps or risks and gates fail
The native review inspects the current WU spec, changed files, acceptance criteria, and delivery
risks. It is intentionally separate from wu:verify, which keeps its existing lifecycle meaning.
Conditional Commands
Section titled “Conditional Commands”Define pattern-triggered commands alongside your standard gates:
| Field | Type | Required | Description |
|---|---|---|---|
trigger_patterns | string[] | Yes | Glob patterns matched against changed files |
command | string | Yes | Shell command to execute when patterns match |
severity | string | No | error (default, blocks gates), warn, or off (skip) |
fresh_on_completion | boolean | No | Rerun after completion reconciliation; defaults to false |
guidance | string | No | Actionable text shown when the command fails |
guidance_ref | string | No | File path whose content is appended to guidance |
How it works:
- When
pnpm gatesorwu:prepruns, changed files are compared against each command’strigger_patternsusing glob matching - Only commands with matching patterns execute — unmatched commands are silently skipped
- If a matching command fails with severity
error, gates fail. With severitywarn, a warning is logged but gates continue - A matching command with
fresh_on_completion: trueruns again duringwu:done, after branch reconciliation and before the landing transaction, even when ordinary gates reuse a validwu:prepcheckpoint - Completion freshness is fail-closed: a failed or unavailable fresh command blocks completion,
including a command whose normal preparation severity is
warn.severity: offstill disables it - Each fresh completion run writes durable flow evidence bound to the WU ID, reconciled commit, diff, command, completion timestamp, and result. Commands without the flag keep normal checkpoint reuse
Registering via the CLI (Constraint-9 compatible):
workspace.yaml must not be edited by hand. Two sanctioned paths exist:
-
gate:conditional(recommended, per-rule) — mirrorsgate:co-change:The
namefield is how--remove/--editaddress a specific entry. It is optional in the underlying schema (existing unnamed entries continue to work) but required for CLI-managed entries. -
config:set --json-value(escape hatch) — writes the whole array verbatim when you need a shape the flags above don’t cover:
Both paths validate against ConditionalCommandConfigSchema and commit atomically
via micro-worktree.
Using Presets
Section titled “Using Presets”For common languages, use a preset to get sensible defaults:
Available presets: node, python, go, rust, dotnet, java, ruby, php
Path-scoped tests.unit execution is preset-aware. When the active preset supports scoped
execution, wu:prep can use the current WU’s tests.unit entries to narrow the test gate. When a
preset does not support path-scoped execution, such as dotnet, LumenFlow falls back to the
configured default test command for that preset instead of attempting a JavaScript-specific runner.
One test plan per gate run
Section titled “One test plan per gate run”The immutable safety-critical-test gate owns the shared test plan. It runs
declared tests.unit paths when they can be scoped safely; otherwise it runs
the configured test_incremental command. The normal test gate reuses that
exact result, so the same test process does not execute twice during one
wu:prep.
lumenflow init writes a safe test_incremental by default when the project exposes an explicit
changed-test script or recognizable Vitest, Jest, Nx, or Turbo evidence. It never copies
test_full into the incremental slot. If the prepared commit and gate inputs remain unchanged,
wu:done reuses the successful prep checkpoint; only fresh completion commands rerun separately.
The planner falls back to test_full when broader coverage is required:
- main-snapshot comparison probes;
--full-testsor--full-coverage;- test-runner configuration changes;
- unsafe or missing scope, unavailable change detection, or untracked code; and
- a missing, blank, or full-equivalent
test_incrementalcommand.
These fallbacks also execute once across the safety and normal test gates. Full CI remains the final whole-repository authority.
Safe prep-evidence reuse (WU-3949)
Section titled “Safe prep-evidence reuse (WU-3949)”Prep-evidence reuse is default-off and is not a substitute for the test-plan safety rules above. Enable the fast path only with both settings:
Only an identical eligible low- or medium-risk plan can reuse passed evidence. Its source, tests,
dependencies, toolchain, main snapshot, policy, authority, applicability, and execution plan must
all match. Missing, malformed, stale, cross-workspace, or mismatched evidence runs gates fresh.
--full-tests, report-all, and high-risk prep run gates fresh; --full-tests also forces full test
execution. Snapshot isolation remains OFF: reuse does not pin or execute an immutable workspace
snapshot.
Uncovered-changed-files cross-check
Section titled “Uncovered-changed-files cross-check”A green scoped run only proves the declared tests.unit paths passed — it
does not by itself prove every changed file was exercised. After the
declared scoped run passes, LumenFlow diffs the branch and checks whether
any changed code file falls outside the declared tests.unit paths. When
every changed file is already covered, nothing further runs and the scoped
run’s cost stays exactly as before.
When changed files remain uncovered, LumenFlow runs one additional, narrowly
scoped vitest related check against just those files: vitest resolves the
module graph and finds any test in the project that imports them, wherever
it lives.
- If that check finds a covering test and it fails, the gate fails
immediately and names the uncovered file(s) so
tests.unitcan be extended to include the covering suite. - If it finds no covering test anywhere (a true coverage gap, not just an
incomplete declaration), the run does not pass silently; it falls through
to the configured
test_incrementalregression flow instead of trusting the declared-only result.
Preset Defaults
Section titled “Preset Defaults”| Preset | Format | Lint | Typecheck | Test |
|---|---|---|---|---|
node | prettier --check . | eslint . | tsc --noEmit | npm test |
python | ruff format --check . | ruff check . | mypy . | pytest |
go | gofmt -l . | golangci-lint | go vet ./... | go test |
rust | cargo fmt --check | cargo clippy | cargo check | cargo test |
dotnet | dotnet format --verify | dotnet build | - | dotnet test |
java | spotless:check | checkstyle | mvn compile | mvn test |
ruby | rubocop | rubocop | - | rspec |
php | php-cs-fixer | phpstan | - | phpunit |
Running Gates
Section titled “Running Gates”Dependency Isolation Preflight
Section titled “Dependency Isolation Preflight”Before gate context, telemetry, or any gate command starts, LumenFlow inspects the active checkout’s
dependency roots and @lumenflow workspace-package links. Each workspace dependency must resolve
inside the active checkout. Package managers may materialize shared content into a checkout-local
virtual store via hardlinks, reflinks, or copies, but a direct workspace link to an external store is
never accepted. Store-looking substrings do not establish trust. Every intermediate scope/package
component is realpath-checked, including directory junctions. Missing/non-directory dependency
roots, missing declared workspace packages, and paths into main or another worktree all fail closed.
The diagnostic includes both the exact contaminated link and its resolved target:
Remove only the listed link, run the configured frozen install inside the active worktree, and then rerun gates. Do not relink to main and do not skip the check: no gate result is trustworthy when module resolution can read another branch.
A worktree gates invocation also requires its own CLI dist. It never falls back to main’s CLI dist,
so a --skip-setup checkout cannot bypass this preflight by bootstrapping the gate runner from a
different checkout.
Main-snapshot comparison probes follow the same boundary. Each temporary probe performs its own frozen install and blocks classification if installation or isolation verification fails.
If any gate fails, wu:prep fails and you fix issues in the worktree before completion.
For migration verification failures, the expected fix is:
- Apply the pending migrations using your project’s normal process
- Re-run the configured verification command manually if needed
- Re-run
pnpm wu:prep --id WU-XXX
Gate Flags
Section titled “Gate Flags”| Flag | Description |
|---|---|
--docs-only | Run only docs-related gates (skip format/lint/typecheck/test) |
--full-tests | Force one full test_full execution instead of scoped or incremental tests |
--full-lint | Run full lint pass instead of scoped lint |
Command Options
Section titled “Command Options”Commands can be strings or objects with options:
Language Examples
Section titled “Language Examples”Node.js / TypeScript
Section titled “Node.js / TypeScript”Python
Section titled “Python”Java / JVM
Section titled “Java / JVM”How Gates Work Under the Hood
Section titled “How Gates Work Under the Hood”Each gate maps to a policy rule in the Software Delivery Pack:
| Gate | Policy ID | Trigger |
|---|---|---|
| Format check | software-delivery.gate.format | on_completion |
| Lint | software-delivery.gate.lint | on_completion |
| Type check | software-delivery.gate.typecheck | on_completion |
| Test | software-delivery.gate.test | on_completion |
When wu:prep or wu:done runs, the kernel evaluates these policies. A deny from any gate makes the completion decision final — the deny-wins invariant applies. The result is recorded in the evidence store for audit.
Gate Behavior
Section titled “Gate Behavior”On Claim
Section titled “On Claim”When you wu:claim:
- Worktree is created
- Gates status is “pending”
During Work
Section titled “During Work”Run pnpm gates frequently:
On Prep and Done
Section titled “On Prep and Done”When you wu:prep:
- Gates run automatically in the worktree
- If any fail, the WU stays
in_progressuntil you fix and rerunwu:prep
When you wu:done:
- The WU merges to main
- The stamp is created
- The worktree is cleaned up
Gate Failures
Section titled “Gate Failures”Fix and retry:
Test gate: capture-integrity failures vs real failures
Section titled “Test gate: capture-integrity failures vs real failures”If the test gate fails and prints [test-evidence:capture-integrity-failure] ..., the captured
process output had no test failure markers at all (for example, spinner/progress-only output with
no per-file results or summary) even though the gate exited non-zero. This means the capture
was broken, not that a specific test failed — rerun the underlying test command directly with a
forced JSON or verbose reporter to get real evidence before treating the run as a regression. See
Gates Reference — test-gate evidence capture-integrity classification
for the full code catalogue.
If it instead prints [test-evidence:partial-output] ... under concurrent gate contention (several
agents running pnpm gates at once), the reason now names an actionable retry step. A duplicated
but agreeing runner summary reconciles automatically; only a genuinely conflicting duplicate still
fails closed. See
Gates Reference — partial-output reconciliation under concurrent-runner contention.
Branch-vs-main comparison artifacts (WU-3836)
Section titled “Branch-vs-main comparison artifacts (WU-3836)”When an immutable gate fails, branch-vs-main attribution pins one main snapshot
and executes a resolved test plan at most once when its full fingerprint matches.
Different applicability (including test, safety-critical-test, and docs-only)
plans never share a result. The comparison remains fail-closed when it is unrunnable
or loses infrastructure. Before the temporary probe worktree is removed, a
bounded JSON artifact is retained at
.lumenflow/artifacts/gate-comparison/<source-sha>-<gate>-<unique-id>.json. It records the
exact command, source SHA, runtime/environment, test name, assertion/error body,
exit or signal, duration, cleanup disposition, bounded-field truncation flags,
original lengths, and SHA-256 digests. At most 100 recent artifacts are retained.
The classification is one of
branch-only-regression, identical-main-failure, unrunnable-comparison,
timeout, or infrastructure-loss.
To reproduce concurrent attribution deterministically, run the filesystem-barrier fixture repeatedly:
Pre-existing, introduced, or undeterminable (WU-3981)
Section titled “Pre-existing, introduced, or undeterminable (WU-3981)”That branch-vs-main attribution settles every blocking failure into one of three outcomes:
- Pre-existing on main — the same failure reproduces on the pinned main snapshot. LumenFlow auto-skips the gate; no action needed.
- Introduced by this branch — main passes, only the branch fails. The gate blocks and cannot be skipped; fix the failure.
- Undeterminable — main also fails, but the failure cannot be matched to the branch’s (an
unrunnable comparison, or a genuinely different failure). The gate stays blocking and prints why
it could not certify pre-existing, followed by a remedy that depends on whether the gate is
skippable at all: a skippable gate gets a named
--skip-gate <name> --reason "<why>" --fix-wu <WU-ID>remedy, with independent verification and the repository owner’s authorization; a non-skippable gate (every gate on the immutable safety-critical deny-set, plustest) gets no skip remedy — fix the failure on this branch, or repair main. There is no automatic skip for either case. See Gates Reference — Three classification outcomes for the full table.
Skip Flags
Section titled “Skip Flags”There is no bare pnpm gates --skip-<gate> flag. Skip a specific, named gate through wu:done
instead — both --reason and --fix-wu are required, and only gates marked skippable: true can
be named this way:
See Constraints — Gates and Named Gate Skips for the gates that are immutable and can never be skipped this way.
Or in configuration:
CI Integration
Section titled “CI Integration”Use the LumenFlow Gates GitHub Action:
The action reads your software_delivery.gates.execution config automatically. See GitHub Action docs for details.
Backwards Compatibility
Section titled “Backwards Compatibility”If no software_delivery.gates.execution config is present, LumenFlow falls back to auto-detecting your project type based on files present and uses preset defaults.
Next Steps
Section titled “Next Steps”- Policy Engine — How gates are evaluated as deny-wins policies
- Evidence Store — How gate results are recorded for audit
- Configuration Reference — Full gates schema
- CLI Reference — All gate commands