Test Series
Aartiq's core modules are verified with an automated Jest test suite. This page reports the real numbers, what is covered, and what is not.
65
Test Suites (65 passing + 0 skipped)
1438
Total Tests (declared blocks)
1426
Passing
12
Skipped · 0 failing
Counts generated on macOS (local) at 2026-10-08T11:55:29.937Z — not typed. CI runs a different platform, so its per-job numbers differ; see the run below.
5/5 jobs success — run #83 (workflow_dispatch, 2026-10-08)Breakdown
Test Suites
| Suite | Tests | What it verifies |
|---|---|---|
| shell-command-tiers | 193 | Shell tier invariants: destructive pattern floors, unknown binaries, URL arguments, blocked-command rejection, Allow Always eligibility, command key normalization |
| sandbox-security | 64 | Fail-closed sandboxing (Seatbelt / bubblewrap / AppContainer + Job Objects), macOS adversarial OS-enforcement (AF_UNIX + signal confinement), command tokenizer, env sanitization |
| shell-approval-defaults | 61 | Startup store contents, denials by risk tier, prompting rules, exact-match Always grants, MCP classifier parity, fail-closed fallbacks, legacy grant migration |
| research-pipeline | 60 | Budget normalization, domain and date parsing, claim extraction, corroboration and dispute logic, recency selection, run events, hard budget enforcement |
| extraction | 58 | Web extractor, DOM parsing, content extraction edge cases (declared test blocks; it.each expands case count) |
| skill-loading | 54 | Dynamic skill loading, validation allowlist, require-path resolution |
| tab-intelligence | 51 | Tab intelligence, domain grouping, smart icons |
| page-scripts-forms | 48 | Stale ref failures, fill and type semantics, per-field form results, gated form submission, element actions, read-only CSS and prose search |
| local-server-auth | 47 | Bind host resolution, token extraction precedence, Host and Origin checks, request gate rejections, loopback bridge sockets, MCP pairing URL flow |
| docs-platform-integration-match-source | 43 | Docs gate: platform action reachability, undelivered scheme handling, IPC channel registration, dropped features and shortcut ids claims pinned to source |
| security-validator | 42 | Blocklist / injection detection / risk classification |
| security-fixes | 41 | Regression suite for applied security fixes |
| dom-engine | 40 | DOM interaction engine, click/fill strategies |
| component-tests | 37 | React component behavior and props |
| directory-allowlist | 36 | Path canonicalization, symlink traversal, read/write separation |
| agent-api-bridge-tools | 35 | Bridge ref resolution and staleness, snapshot options, page reading tools, navigation stamp reset, search delegation, tool descriptors, fill versus submit gating |
| web-search-service | 30 | Provider key resolution and fallback order, result count clamps, keyed versus scraped labeling, news provider routing, result parsing, key-safe provider info |
| webauthn-service | 26 | WebAuthn / FIDO2 challenge-response flow |
| apple-intelligence-genmoji | 24 | Genmoji export identity, single-command spawn with JSON prompt transport, off-macOS refusal, IPC channel exposure, helper OS floors, docs command coverage |
| research-progress-plumbing | 24 | Job reducer transitions, progress clamping and step handling, terminal failure stages, event filtering, main and preload wiring, superseded job filtering |
| snapshot-ref-binding | 24 | axId stamping and marker cleanup, ref lifecycle and staleness, structural node refusal, single-page search filters, limits and miss behavior |
| file-paths | 22 | URL versus file path detection, path wrapping preprocessing, tokenizer classification, code block protection, idempotence, rendered link and chip output |
| approval-ticket-security | 21 | Ticket-based approval + capability-controller regression (audit findings) |
| docs-deep-links-match-source | 19 | Docs gate: deep-link command statuses, dead command reachability, documented parameters, retired route and mechanism claims pinned to source |
| docs-sync-match-source | 17 | Docs gate: Cloud Sync trust levels, discovery mechanism, size and duration limits, feature capabilities and encryption scoping claims pinned to source |
| automation | 16 | OS automation layer (click / scroll / app launch) |
| dom-handlers | 16 | Browser DOM IPC handlers |
| sensitive-paths-deny | 15 | Sensitive-path denial over allowlisted homes and workspaces, symlink realpath resolution, read-only directory defaults, broad grant narrowing, sandbox deny blocks |
| sync-auth-tokens | 15 | Pairing token issuance, listener authentication and lifetime checks, refresh device binding, unpair revocation, pairing lockout, HTTP token and Host validation |
| linux-bwrap-sandbox | 147 skipped | bubblewrap arg generation, namespace/unshare flags (pid/net/ipc/uts/user/cgroup), capability pre-flight fail-closed, Linux runtime enforcement |
| cloud-sync-error-surface | 13 | Rejected cloud write logging and error emission across source and compiled twins, failure forwarding to renderer, status line surfacing, success silence |
| docs-listener-claims-match-source | 13 | Docs gate: listener binding, token coverage, approval tier tables, enforcement layer and tool count claims pinned to source |
| windows-job-sandbox | 135 skipped | Windows AppContainer JS contract + runtime matrix (suspended AppContainer start, OS-enforced ACL allowlist, verified job, grandchild containment, secret isolation, KILL_ON_JOB_CLOSE) |
| allow-always-lifetime | 12 | Grant lifetime policy constant, grant record persistence and revocation, gate lifetime enforcement and sweeping, pre-lifetime grant migration |
| permission-store-system-root | 12 | System root rejection in allowed directories, trailing slash handling, exact-match path boundary rule, normal directory acceptance, rejection auditing |
| approval-gate-concurrency | 11 | Ticket single redemption under concurrency, sequential replay refusal, input hash burning, scope mismatch and status reporting, live redemption path |
| wifi-sync-upgrade | 11 | WebSocket upgrade Origin and Host gating, DNS rebinding and foreign origin refusal over real sockets, access token gating of unpair |
| network-listener-hardening | 10 | Loopback-only binding for bridge and services, wildcard-free listen calls, distinct default ports across agent API and native bridge, routing targets |
| remote-shell-approval | 10 | Remote origin shell approval registration policy, QR and PIN ticket flow, invalid and tampered input denial, single use redemption, sandbox fallback |
| native-bridge-permissions | 9 | Permission grant, revoke and read routes, immediate gate visibility, lost-update resistance, invalid level fail-closed handling, audit trail |
| session-token | 8 | Per-listener token file creation and mode, cross-restart persistence, listener separation, corrupt file replacement, unwritable home fallback, rotation |
| citation-links | 7 | Citation bracket to markdown link normalization, verbatim URL copying, idempotence, rendered href correctness, plain text and bracket preservation |
| docs-shortcuts-match-source | 7 | Docs gate: published shortcut table, registered accelerator parity, duplicate and unbound entries, scoped key descriptions and conflict surfacing claims pinned to source |
| native-approval-biometrics | 7 | Dialog label honesty, deny handling, biometric flag enforcement with fail-closed verification and unsupported platforms, per session caching |
| extensions.crx-verifier | 6 | CRX3 signature enforcement — Chromium's format: bounds-checked header parse, crx_id ↔ signing-key binding, and signatures over the signed header plus the zip archive (6 tests, runs everywhere) |
| guardrails.origin-guard | 6 | Origin trust levels and verb-scoped permission gates |
| guardrails.prompt-injection | 6 | Prompt-injection detection guards |
| licence-audit-rename | 6 | Root licence text, manifest and installer licence declaration, persisted audit file rename migration without clobbering, chat export default naming |
| linux-ipc-registration | 6 | Linux bridge channel registration uniqueness against main, preload invoked channel ownership, platform guard error objects on non-Linux |
| pairing-auth | 6 | Deterministic master key signature computation and verification, wrong key and tampered signature rejection, fail-closed missing inputs, replay window |
| snapshot | 6 | Snapshot / accessibility-tree reference stability |
| sync-handlers-window | 6 | Live window resolution for prompt delivery, honest failure when windows are gone, destroyed window safety, captured window fallback, windowless status |
| wifi-sync-pairing | 6 | Pairing code stability across restarts, unknown device code checks, paired device id recognition with trusted legacy ids, untrusted legacy rejection |
| agent-api.registry | 5 | Agent API tool registry and call routing |
| agent.tab-lock | 5 | Per-tab locking between concurrent agents |
| benchmark-smoke | 5 | Benchmark runner completion, expected benchmark presence and timing shape, environment capture, KDF work floors, unknown only flag rejection |
| local-server-auth-lockout | 5 | Remote address failure counting and lockout, lock clearing on success, shared address family counters, loopback exemption, socketless local handling |
| agent.trust | 4 | Agent trust-level assignment and scoping |
| ai-command-parser-plan | 4 | PLAN and THINK payload field extraction, bracket form parsing, reasoning note capture so plan and thinking text stays populated |
| home-intelligence | 4 | Home intelligence logic |
| theme | 4 | Theme and UI-mode switching |
| approval-gate | 3 | Approval ticket race prevention, one-time consumption, sequential replay refusal, mismatched input burning |
| autofill.vault | 3 | Encrypted autofill vault (AES-GCM, passphrase-derived key) |
| extensions.permission-analyzer | 3 | Chrome extension manifest permission analysis |
| markdown-render | 3 | Currency text kept literal in chat markdown, dollar amounts preserved across sentences, double dollar math still rendered |
Continuous Integration
Latest CI Run
All five jobs were green on the run above — four Jest jobs (full suite, macOS Seatbelt, Linux bubblewrap, Windows AppContainer) plus a typecheck job (tsc --noEmit) that reports no test counts; per-job results live on the testing page. Dispatch inputs can reduce the Jest jobs to 3 (skip-full-suite) or 1 (windows-test-pattern), so this is a default-dispatch count rather than an invariant. The run below is #83 (workflow_dispatch, 2026-10-08, conclusion: success).
| Job | OS | Passed | Skipped | Failed | Declared |
|---|---|---|---|---|---|
| Run Jest (aartiq-browser) | ubuntu-latest | 1412 | 26 | 0 | 1438 |
| Run Jest (Windows AppContainer sandbox runtime) | windows-latest | 61 | 30 | 0 | 91 |
| Run Jest (macOS Seatbelt sandbox runtime) | macos-latest | 105 | 0 | 0 | 105 |
| Run Jest (Linux bubblewrap sandbox runtime) | ubuntu-latest | 57 | 21 | 0 | 78 |
| All 5 jobs | 1635 | 77 | 0 | 1712 | |
Why tests are skipped
Breakdown from the generated test facts (2026-10-08): 12 skipped in total.
| Platform-skipped | 12 |
Coverage
What's Covered
- Fail-closed by construction: every sandbox setup, validation, or policy failure returns a structured SANDBOX_* error and the command is never silently run unsandboxed — there is no automatic fallback path
- CRX3 extension packages — the verifier suite (6 tests) parses Chromium's header format with bounds-checked varints (malformed input fails closed instead of hanging), binds crx_id to the signing key, checks signatures over the signed header plus the zip archive, and rejects ZIP end-of-central-directory tokens inside the header; a second suite (tests/crx-url-binding.test.js, 17 tests) pins the install-path binding — the Web Store URL's declared id must equal the verified package's crx_id, asserted against the real verifier with a real signed package; both run in the full-suite job and locally on every platform (publisher-key allowlisting is not implemented — see Known Limits)
- macOS Seatbelt — real OS enforcement: writing outside the directory allowlist is denied by the kernel and the file is verified absent; reading a secret outside the allowlist is denied; /tmp is writable; an IP network bind is denied; an AF_UNIX socket bind is denied; signalling a host process is denied while self-signal works; reading/writing through a symlink that escapes the allowlist is denied; a child process spawned by the target is still contained
- Linux bubblewrap — closed-by-default namespaces (pid/net/ipc/uts/user/cgroup + new session), correct --bind (write) vs --ro-bind (read-only) mapping, network denied by default, and fail-closed when bwrap is missing OR present-but-incapable of creating the required namespaces (the capability pre-flight)
- Windows AppContainer — policy fail-closed (missing runner, invalid allowlist, network-allowlist requests), result parsing, explicit isolation flags ({ filesystem:true, network:true, process:true }), plus a runtime matrix proving suspended AppContainer start + OS-enforced ACL allowlist + verified job assignment + grandchild containment + secret isolation + KILL_ON_JOB_CLOSE
- Explicit isolation contract — every result carries { filesystem, network, process }; macOS/Linux/Windows all report all-true when their platform sandbox is active, and any setup failure or unsandboxed run reports all-false
- Directory allowlist — fs.realpath() canonicalization, ../ traversal, symlink escape, read-only vs read-write separation, and invalid/missing-path rejection (never silently skipped)
- Command execution — the tokenizer preserves quoted arguments verbatim, separates direct execution from explicit shell mode, and never reconstructs a command via a string-joined sh -c; it is documented as a classifier, not a security parser
- Environment sanitization — API keys, tokens, and secrets are stripped from every sandboxed process; only an allowlisted set of non-credential variables passes through
- Security regressions — a 41-test regression suite re-verifying each applied security fix, plus approval-ticket tests that lock in the audit remediation
Coverage Boundaries
Known Limits
What this suite does NOT prove
We would rather state these limits plainly than overstate coverage.
- Runtime enforcement tests only EXECUTE on their own OS. All five jobs were green on the run above — four Jest jobs (full suite, macOS Seatbelt, Linux bubblewrap, Windows AppContainer) plus a typecheck job (tsc --noEmit) that reports no test counts; per-job results live on the testing page. Dispatch inputs can reduce the Jest jobs to 3 (skip-full-suite) or 1 (windows-test-pattern), so this is a default-dispatch count rather than an invariant. Latest run [#83](https://github.com/Latestinssan/Aartiq/actions/runs/37772437527) (workflow_dispatch, 2026-10-08) reported: Run Jest (aartiq-browser) on ubuntu-latest 1412 passed / 26 skipped / 0 failed of 1438; Run Jest (Windows AppContainer sandbox runtime) on windows-latest 61 passed / 30 skipped / 0 failed of 91; Run Jest (macOS Seatbelt sandbox runtime) on macos-latest 105 passed / 0 skipped / 0 failed of 105; Run Jest (Linux bubblewrap sandbox runtime) on ubuntu-latest 57 passed / 21 skipped / 0 failed of 78. Suspended AppContainer start, OS-enforced ACL allowlist, verified job assignment, grandchild containment, secret isolation, and KILL_ON_JOB_CLOSE all return verified sandbox results.
- macOS Seatbelt OS-enforcement tests execute only on macOS; they pass on this machine and run in CI on macos-latest. The profile-generation and fail-closed config paths are asserted on every platform.
- These are unit and integration tests for core modules. They do NOT cover the full Electron UI, installers, MSIX/MSI packaging, or complete end-to-end user flows.
- A sandbox confines what code can do; it is not a proof that the AI's decisions are safe, nor a substitute for least-privilege OS accounts, patched dependencies, or simply not running untrusted code. See the security page's 'What this does NOT guarantee'.
- Counts above are declared it()/test() blocks as of 2026-10-08 on macOS (local); suites using it.each expand into more executed cases. Run npx jest (Node 24+, some deps are ESM) or check the CI run for jest.yml for exact pass/skip/fail numbers.
Known limits in the product
The shared list kept in the source of truth — product-wide, not just this suite. The repository README points here instead of duplicating it.
- Runtime sandbox tests execute only on their own OS. There is no single job that exercises Seatbelt, bubblewrap, and AppContainer at once.
- OS-automation tests skip wherever their native tooling or a display is absent (xdotool/xte on Linux, cliclick on macOS). CI installs the tooling and runs them under Xvfb, but a machine without them still skips the suite.
- The CRX3 verifier enforces Chromium's checks — the crx_id must be derived from a key whose signature also verifies, and the signature covers the signed header and the zip archive — and installFromWebStore binds the install to the id the download URL declares (…x=id%3D<32-char id>…): the verified package's crx_id must equal it, so a URL promising one extension can only install that exact extension, never a validly signed different one. What it still does not implement is Chrome's publisher-key allowlisting — an attacker who controls the URL can name a fresh id of their own key, which stays equivalent to sideloading a new extension rather than hijacking an existing one.
- SecurityValidator.js does not guarantee that non-blocked commands are safe — it is a fast first-pass reject layer.
- Visual extraction reduces the DOM-based prompt-injection surface. It does not prevent prompt injection, and it cannot give semantic immunity against instructions rendered into the viewport.
- Seatbelt profiles start from (deny default): an operation class the profile does not explicitly grant is denied, unknown-unknowns included. Three plumbing grants keep real commands working — sysctl-read (read-only introspection), mach-lookup (Mach service-name discovery, so node/python/shell keep working), and file-ioctl (bounded by the file allowlist) — while everything else the old default-allow baseline left open (foreign process-info, user preferences, host-level mach ops) is now denied.
- Apple Events cannot be filtered by the current sandbox-exec — probing it rejects a deny rule with "unbound variable: apple-events", so the operation is not exposed at all — and a sandboxed command could still ask another app to act on its behalf.
- The WiFi sync server (3004) binds every network interface on purpose — the phone reaches it over the LAN — so the LAN exposure itself is the limit: the upgrade refuses foreign Origins and Host headers that do not name this machine, every sync action (unpair included) requires the device's short-lived access token, and AARTIQ_WIFI_SYNC_HOST narrows the bind when that exposure is not wanted. See network.servers.
- The session tokens for the MCP bridge, the Agent API and the native bridge persist in mode-0600 files in your home directory (~/.aartiq-mcp-token, ~/.aartiq-agent-token, ~/.aartiq-token), so a client configured once keeps working across restarts, and each listener now also accepts per-client credentials — mint one with the primary token (POST /clients), revoke just that one (POST /clients/revoke), peers and primary untouched — but remote mode is still not a finished design: there is no pairing UI, nothing writes a credential to a remote client, and the binds are not operator-named.
- "Allow Always" is keyed on the full normalised command line, which is narrower than before but is still text matching — it records what the command says, not what it will do — and every grant now expires after 30 days, swept with an audit-log entry, so the dialog asks again. See aartiq-browser/docs-audit/issues/allow-always-granularity.md.
- An Allow Always grant requires a binary that appears in the classifier's table. One that does not — including anything we have never seen — is offered Allow Once only, because a grant that repeats a command nobody can describe is a promise about behaviour rather than about the text. Local writes such as cp, mv, mkdir and touch are in the table and keep exact-match Always grants, which expire after 30 days.
Reproduce
Run the Tests
Install
cd aartiq-browser
npm installRun full suite
npx jestRun one suite
npx jest tests/sandbox-security.test.js