Aartiq
Aartiq™
Download
Testing & Test Series

Test Series

Aartiq's core modules are verified with an automated Jest test suite. This page reports the real numbers, what is covered, and what is not.

65

Test Suites (65 passing + 0 skipped)

1438

Total Tests (declared blocks)

1426

Passing

12

Skipped · 0 failing

Counts generated on macOS (local) at 2026-10-08T11:55:29.937Z — not typed. CI runs a different platform, so its per-job numbers differ; see the run below.

5/5 jobs success — run #83 (workflow_dispatch, 2026-10-08)

Breakdown

Test Suites

SuiteTests
shell-command-tiers193
sandbox-security64
shell-approval-defaults61
research-pipeline60
extraction58
skill-loading54
tab-intelligence51
page-scripts-forms48
local-server-auth47
docs-platform-integration-match-source43
security-validator42
security-fixes41
dom-engine40
component-tests37
directory-allowlist36
agent-api-bridge-tools35
web-search-service30
webauthn-service26
apple-intelligence-genmoji24
research-progress-plumbing24
snapshot-ref-binding24
file-paths22
approval-ticket-security21
docs-deep-links-match-source19
docs-sync-match-source17
automation16
dom-handlers16
sensitive-paths-deny15
sync-auth-tokens15
linux-bwrap-sandbox147 skipped
cloud-sync-error-surface13
docs-listener-claims-match-source13
windows-job-sandbox135 skipped
allow-always-lifetime12
permission-store-system-root12
approval-gate-concurrency11
wifi-sync-upgrade11
network-listener-hardening10
remote-shell-approval10
native-bridge-permissions9
session-token8
citation-links7
docs-shortcuts-match-source7
native-approval-biometrics7
extensions.crx-verifier6
guardrails.origin-guard6
guardrails.prompt-injection6
licence-audit-rename6
linux-ipc-registration6
pairing-auth6
snapshot6
sync-handlers-window6
wifi-sync-pairing6
agent-api.registry5
agent.tab-lock5
benchmark-smoke5
local-server-auth-lockout5
agent.trust4
ai-command-parser-plan4
home-intelligence4
theme4
approval-gate3
autofill.vault3
extensions.permission-analyzer3
markdown-render3

Continuous Integration

Latest CI Run

All five jobs were green on the run above — four Jest jobs (full suite, macOS Seatbelt, Linux bubblewrap, Windows AppContainer) plus a typecheck job (tsc --noEmit) that reports no test counts; per-job results live on the testing page. Dispatch inputs can reduce the Jest jobs to 3 (skip-full-suite) or 1 (windows-test-pattern), so this is a default-dispatch count rather than an invariant. The run below is #83 (workflow_dispatch, 2026-10-08, conclusion: success).

JobOSPassedSkippedFailedDeclared
Run Jest (aartiq-browser)ubuntu-latest14122601438
Run Jest (Windows AppContainer sandbox runtime)windows-latest6130091
Run Jest (macOS Seatbelt sandbox runtime)macos-latest10500105
Run Jest (Linux bubblewrap sandbox runtime)ubuntu-latest5721078
All 5 jobs16357701712

Why tests are skipped

Breakdown from the generated test facts (2026-10-08): 12 skipped in total.

Platform-skipped12

Coverage

What's Covered

  • Fail-closed by construction: every sandbox setup, validation, or policy failure returns a structured SANDBOX_* error and the command is never silently run unsandboxed — there is no automatic fallback path
  • CRX3 extension packages — the verifier suite (6 tests) parses Chromium's header format with bounds-checked varints (malformed input fails closed instead of hanging), binds crx_id to the signing key, checks signatures over the signed header plus the zip archive, and rejects ZIP end-of-central-directory tokens inside the header; a second suite (tests/crx-url-binding.test.js, 17 tests) pins the install-path binding — the Web Store URL's declared id must equal the verified package's crx_id, asserted against the real verifier with a real signed package; both run in the full-suite job and locally on every platform (publisher-key allowlisting is not implemented — see Known Limits)
  • macOS Seatbelt — real OS enforcement: writing outside the directory allowlist is denied by the kernel and the file is verified absent; reading a secret outside the allowlist is denied; /tmp is writable; an IP network bind is denied; an AF_UNIX socket bind is denied; signalling a host process is denied while self-signal works; reading/writing through a symlink that escapes the allowlist is denied; a child process spawned by the target is still contained
  • Linux bubblewrap — closed-by-default namespaces (pid/net/ipc/uts/user/cgroup + new session), correct --bind (write) vs --ro-bind (read-only) mapping, network denied by default, and fail-closed when bwrap is missing OR present-but-incapable of creating the required namespaces (the capability pre-flight)
  • Windows AppContainer — policy fail-closed (missing runner, invalid allowlist, network-allowlist requests), result parsing, explicit isolation flags ({ filesystem:true, network:true, process:true }), plus a runtime matrix proving suspended AppContainer start + OS-enforced ACL allowlist + verified job assignment + grandchild containment + secret isolation + KILL_ON_JOB_CLOSE
  • Explicit isolation contract — every result carries { filesystem, network, process }; macOS/Linux/Windows all report all-true when their platform sandbox is active, and any setup failure or unsandboxed run reports all-false
  • Directory allowlist — fs.realpath() canonicalization, ../ traversal, symlink escape, read-only vs read-write separation, and invalid/missing-path rejection (never silently skipped)
  • Command execution — the tokenizer preserves quoted arguments verbatim, separates direct execution from explicit shell mode, and never reconstructs a command via a string-joined sh -c; it is documented as a classifier, not a security parser
  • Environment sanitization — API keys, tokens, and secrets are stripped from every sandboxed process; only an allowlisted set of non-credential variables passes through
  • Security regressions — a 41-test regression suite re-verifying each applied security fix, plus approval-ticket tests that lock in the audit remediation

Coverage Boundaries

Known Limits

What this suite does NOT prove

We would rather state these limits plainly than overstate coverage.

  • Runtime enforcement tests only EXECUTE on their own OS. All five jobs were green on the run above — four Jest jobs (full suite, macOS Seatbelt, Linux bubblewrap, Windows AppContainer) plus a typecheck job (tsc --noEmit) that reports no test counts; per-job results live on the testing page. Dispatch inputs can reduce the Jest jobs to 3 (skip-full-suite) or 1 (windows-test-pattern), so this is a default-dispatch count rather than an invariant. Latest run [#83](https://github.com/Latestinssan/Aartiq/actions/runs/37772437527) (workflow_dispatch, 2026-10-08) reported: Run Jest (aartiq-browser) on ubuntu-latest 1412 passed / 26 skipped / 0 failed of 1438; Run Jest (Windows AppContainer sandbox runtime) on windows-latest 61 passed / 30 skipped / 0 failed of 91; Run Jest (macOS Seatbelt sandbox runtime) on macos-latest 105 passed / 0 skipped / 0 failed of 105; Run Jest (Linux bubblewrap sandbox runtime) on ubuntu-latest 57 passed / 21 skipped / 0 failed of 78. Suspended AppContainer start, OS-enforced ACL allowlist, verified job assignment, grandchild containment, secret isolation, and KILL_ON_JOB_CLOSE all return verified sandbox results.
  • macOS Seatbelt OS-enforcement tests execute only on macOS; they pass on this machine and run in CI on macos-latest. The profile-generation and fail-closed config paths are asserted on every platform.
  • These are unit and integration tests for core modules. They do NOT cover the full Electron UI, installers, MSIX/MSI packaging, or complete end-to-end user flows.
  • A sandbox confines what code can do; it is not a proof that the AI's decisions are safe, nor a substitute for least-privilege OS accounts, patched dependencies, or simply not running untrusted code. See the security page's 'What this does NOT guarantee'.
  • Counts above are declared it()/test() blocks as of 2026-10-08 on macOS (local); suites using it.each expand into more executed cases. Run npx jest (Node 24+, some deps are ESM) or check the CI run for jest.yml for exact pass/skip/fail numbers.

Known limits in the product

The shared list kept in the source of truth — product-wide, not just this suite. The repository README points here instead of duplicating it.

  • Runtime sandbox tests execute only on their own OS. There is no single job that exercises Seatbelt, bubblewrap, and AppContainer at once.
  • OS-automation tests skip wherever their native tooling or a display is absent (xdotool/xte on Linux, cliclick on macOS). CI installs the tooling and runs them under Xvfb, but a machine without them still skips the suite.
  • The CRX3 verifier enforces Chromium's checks — the crx_id must be derived from a key whose signature also verifies, and the signature covers the signed header and the zip archive — and installFromWebStore binds the install to the id the download URL declares (…x=id%3D<32-char id>…): the verified package's crx_id must equal it, so a URL promising one extension can only install that exact extension, never a validly signed different one. What it still does not implement is Chrome's publisher-key allowlisting — an attacker who controls the URL can name a fresh id of their own key, which stays equivalent to sideloading a new extension rather than hijacking an existing one.
  • SecurityValidator.js does not guarantee that non-blocked commands are safe — it is a fast first-pass reject layer.
  • Visual extraction reduces the DOM-based prompt-injection surface. It does not prevent prompt injection, and it cannot give semantic immunity against instructions rendered into the viewport.
  • Seatbelt profiles start from (deny default): an operation class the profile does not explicitly grant is denied, unknown-unknowns included. Three plumbing grants keep real commands working — sysctl-read (read-only introspection), mach-lookup (Mach service-name discovery, so node/python/shell keep working), and file-ioctl (bounded by the file allowlist) — while everything else the old default-allow baseline left open (foreign process-info, user preferences, host-level mach ops) is now denied.
  • Apple Events cannot be filtered by the current sandbox-exec — probing it rejects a deny rule with "unbound variable: apple-events", so the operation is not exposed at all — and a sandboxed command could still ask another app to act on its behalf.
  • The WiFi sync server (3004) binds every network interface on purpose — the phone reaches it over the LAN — so the LAN exposure itself is the limit: the upgrade refuses foreign Origins and Host headers that do not name this machine, every sync action (unpair included) requires the device's short-lived access token, and AARTIQ_WIFI_SYNC_HOST narrows the bind when that exposure is not wanted. See network.servers.
  • The session tokens for the MCP bridge, the Agent API and the native bridge persist in mode-0600 files in your home directory (~/.aartiq-mcp-token, ~/.aartiq-agent-token, ~/.aartiq-token), so a client configured once keeps working across restarts, and each listener now also accepts per-client credentials — mint one with the primary token (POST /clients), revoke just that one (POST /clients/revoke), peers and primary untouched — but remote mode is still not a finished design: there is no pairing UI, nothing writes a credential to a remote client, and the binds are not operator-named.
  • "Allow Always" is keyed on the full normalised command line, which is narrower than before but is still text matching — it records what the command says, not what it will do — and every grant now expires after 30 days, swept with an audit-log entry, so the dialog asks again. See aartiq-browser/docs-audit/issues/allow-always-granularity.md.
  • An Allow Always grant requires a binary that appears in the classifier's table. One that does not — including anything we have never seen — is offered Allow Once only, because a grant that repeats a command nobody can describe is a promise about behaviour rather than about the text. Local writes such as cp, mv, mkdir and touch are in the table and keep exact-match Always grants, which expire after 30 days.

Reproduce

Run the Tests

Install

cd aartiq-browser
npm install

Run full suite

npx jest

Run one suite

npx jest tests/sandbox-security.test.js