KEEN documentationOpen KEEN

KEEN developer’s guide

This guide explains the code paths behind the application’s framework, evidence, risk and assurance workflows. It accompanies the adopter’s guide and system administrator’s guide. Paths are relative to the KEEN repository unless another project is named.

1. Architecture and repository map

KEEN uses a FastAPI application, SQLAlchemy models and Alembic migrations over PostgreSQL. Celery worker and beat use Valkey as broker/backend; Valkey also stores sessions, authorisation versions, caches and notification messages. The static UI uses HTML, ES modules, Bootstrap and D3, with shared MOSP design-system helpers.

services/api/app/main.py constructs the application, middleware and startup hooks. api/routes/__init__.py assembles domain routers. db/models.py contains the ORM models; db/session.py provides the engine, SessionLocal and the request database dependency. Schema history lives in services/api/alembic/versions, outside the app package.

core/config.py defines environment settings. core/managed_configuration.py resolves database-managed configuration. security contains authentication, sessions, OIDC, roles, permissions, CSRF, rich-text sanitisation and redaction. ingest contains source adapters and common persistence/SSRF helpers. mapping contains collection identity and rule evaluation. worker contains the Celery application and tasks. services contains changelogs, statistics, scheduling and domain helpers. storage/s3.py provides artefact storage behaviour.

services/ui/public contains HTML pages and their pages/*.js controllers. public/app.js wraps and re-exports the MOSP helpers. The UI Docker build expects mosp-design-system beside keen; keep the design-system version compatible with the application.

2. Follow a request through the application

The external HTTPS proxy forwards to keen-ui. nginx serves static files and forwards /api/ to keen-api with that prefix removed. A browser call to /api/v1/events therefore reaches /v1/events in FastAPI. The OpenAPI server prefix reflects this public path.

Middleware establishes request identity and applies permission and CSRF checks before the handler. Starlette runs inbound middleware in reverse registration order: inspect the registration order when changing it, rather than assuming the order of conceptual concerns. Authentication must populate the request’s user before checks that rely on it.

Authentication uses a short-lived database session and detaches the user before passing the request onward. Handlers use their own get_db dependency. Do not attach a middleware ORM instance to a second session or leave a transaction open across slow external calls.

Pydantic payload models validate shape and restrict writable fields. Domain handlers validate relationships and permissions before writing. Use SQLAlchemy or bound parameters, sanitise rich text on the server, and record semantic changes through the shared entity-changelog service. Request audit logs and entity changelogs serve different purposes.

The request-size middleware checks the declared length. Keep proxy body-size enforcement too; a header check is not a general streaming-body quota. Do not describe an nginx page-hiding rule as API authorisation.

3. Identity, sessions and authorisation

Local authentication, supported identity-provider flows and optional trusted-header authentication resolve to application users. Normal browser sessions are stored in Valkey and referenced by an HttpOnly cookie. Active-account checks and effective permissions remain server-side decisions.

Roles are admin, normal or inherit. Inherited roles are resolved from groups; permissions combine direct and group grants. Sessions cache an authorisation snapshot and a version. Access-changing operations must bump the affected authorisation versions so later requests recompute permissions. A cache miss must not become an allow decision.

The middleware enforces namespace and write permissions, with an admin-only fallback for unrecognised writes. Domain handlers also enforce precise permissions and relationship rules. A new route needs both a deliberate middleware policy and a route-level check; hiding its button is insufficient.

OIDC uses state, nonce and PKCE, a single-use login-state record, token exchange and signature/issuer/audience/expiry validation. Identity links use issuer and subject. General deployment settings control auto-link and provisioning behaviour; do not assume every IdP mode has the same policy. Do not authorise an account simply because a supplied email looks familiar.

The dedicated-hosting integration adds security/hosted.py and bootstrap_hosted.py. Hosted mode requires the configured issuer and owner subject, restricts existing sessions to the owner and disables alternative local/header sign-in paths. Bootstrap creates or binds the owner deliberately. Tenant identity checks apply before normal OIDC user resolution and remain necessary even with an Auth0-side action.

Cookie-based writes use the CSRF cookie/header check. Hosted mode additionally checks the exact origin, including WebSocket origin checks. Webhooks use their own authentication and must remain separately protected. Preserve host-only cookies and do not broaden them to a parent domain shared by customers.

4. Frameworks are application data

api/routes/frameworks.py exposes framework listing/upsert and the administrator node editor. The UI controller is pages/framework-editor.js embedded in admin.html. It manages framework metadata and Clause, Control and Custom nodes with descriptions, upstream URLs, ordering, scope and parent relationships.

ControlItem identities include framework, kind and reference. Clause nodes also maintain first-class FrameworkClause records used by clause views and ControlClauseLink relationships. Keep those representations consistent when modifying the editor. Parents must belong to the appropriate framework/category and form a valid hierarchy.

The editor treats framework slugs and node kind/reference as stable. Deletion checks children, clause links and mappings rather than silently cascading through evidence. Extend dependency handling deliberately when adding another relationship type.

Migration 0074_seed_frameworks seeds the bundled catalogues, including KEEN-AF:1.0. Bundled JSON is migration/build input, not the everyday editing workflow. Seed logic preserves existing populated administrator content where specified. Do not fetch mutable upstream standard text during a migration or assume a catalogue grants rights to licensed normative wording.

5. Evidence definitions and managed configuration

An evidence definition combines a source adapter, collection input, matching conditions and framework targets. The main API implementation is api/routes/managed_configurations.py; the editor is pages/evidence-config.js within Admin → Evidence sources & rules.

ManagedConfiguration stores named documents and their versions. ManagedConfigurationRevision stores prior documents. core/managed_configuration.py reads a database document first and falls back to a shipped input only when no database document exists. The migration chain initialises these documents so migrated installations use PostgreSQL as the working source of configuration.

PUT /v1/admin/evidence-definitions/{rule_id} validates the adapter, input section, stable collection key, rule identity and targets. For a collection-specific definition it locks configuration documents in a consistent order, checks both submitted versions and commits the input and rule revisions together. A 409 tells the caller to reload and reconcile. Preserve this transaction boundary; publishing a query without its matching rule is a partial update.

Source-wide definitions use the selected source without a collector condition. Collection-specific definitions obtain their collector identity through shared mapping helpers. Source identity and collection identity are distinct: several queries can emit events with the same adapter source.

The source catalogue exposes enablement and observed activity without sending credentials to the browser. The BookStack catalogue makes server-side requests using configured credentials. The JSON panels are for collection and adapter settings, with secret-like content rejected; actual credentials remain environment settings.

Rule parsing, matching and preview

mapping/rules.py parses the managed rules document and evaluates events. Conditions include exact source/system/actor/action/outcome/severity/label values, supported regex fields, collection identity and BookStack-specific selectors. Filled conditions combine as AND. Label matching also considers values extracted from supported payload structures.

Rules contain IDs, enabled state, conditions, targets and confidence. Targets identify a framework and reference; one rule can map to several frameworks. Admin validation checks regex and target validity, while the common parser is shared with runtime evaluation. Keep preview and ingestion semantics aligned when adding a field.

The editor supports sampled preview and an explicit predefinition path when no events exist. Preview should be treated as a diagnostic operation, not an authorisation mechanism or proof of compliance. Suggested controls remain proposals for an administrator to accept.

6. Ingestion, artefacts and historical mapping

Pull adapters and incoming webhooks normalise source data into events. ingest/common.py handles shared persistence, deduplication, redaction and artefacts. Stable source/external identifiers prevent duplicate events. Collection provenance must survive normalisation so a collection-specific rule can identify its own evidence.

Keep outbound requests behind the shared URL safety checks, TLS validation and restricted redirect behaviour. Administrator-controlled URLs can still reach dangerous network locations if SSRF protection is weakened. Webhooks must preserve secret validation, replay controls and fail-closed rate limits.

Artefacts have stored references and integrity metadata. In the hosted integration, S3 writes are conditional, and stored references include object version information for retrieval. The runtime uses its tenant role and must not accept a reference to an arbitrary bucket. Do not discard version information when generating a download or presigned URL.

Saving a rule can queue historical evaluation. Rule-backfill records persist progress; worker tasks process jobs and beat schedules recovery. The job has its own lifecycle after the definition transaction, so callers must not equate a saved configuration with a finished backfill.

The broader remap action is additive and can include already-mapped events. Stale-rule-mapping preview/prune is a distinct path: it limits the review batch, rechecks the current rule set and preserves manual/imported mappings. Test rule tightening, deletion, concurrent edits and cancellation separately from straightforward creation.

Control-evidence statistics and caches are maintained separately from event storage. A discrepancy can arise from stale derived data rather than missing evidence. Use the statistics service instead of inventing another counting model in a new endpoint.

7. Risk register, library and KEEN Assurance Framework

api/routes/risks.py handles scenarios and control links. api/routes/risk_register.py adds register views, ratings, thresholds, CSV exchange and reusable templates. Risk is the shared scenario entity. The editor, register and heatmap use the same inherent and residual likelihood/impact fields, each constrained to 1–5, with scores calculated as their product.

RiskLibraryEntry stores reusable scenario text, CIA associations, treatment guidance and suggested assessment data. The library is not a parallel customer risk register. Creating from a template produces a reviewable risk; subsequent template changes do not automatically rewrite that assessment.

The frontend loads the chosen template, example ratings and suggested asset. Risk creation receives library_template_id. The server retrieves suggested KEEN-AF control references from the stored template, filters them against remaining controls and writes those links under KEEN-AF:1.0. This also works when the active framework is different, preserving separate framework-specific links.

Migrations 0076_seed_assurance_risk_library, 0077_unified_risks_assets and 0078_library_control_mappings supply the library, unified risk/asset behaviour and template mappings. Use their actual revision identifiers when inspecting migration history; filenames and revision IDs are not always identical.

CSV import creates new scenarios and validates the complete file before committing. It has size/row limits and validates ownership, ratings and control references. Export protects spreadsheet cells against formula interpretation. Preserve whole-file rejection and formula-safe output when changing the schema.

8. Documents, people and assurance records

Native documents support folders, types, tags, rich-text content, immutable content revisions, hashes and comments. The UI’s update includes expected_content_version to detect concurrent content edits. Preserve server-side sanitisation and version checks when extending the document editor.

People and personnel-assurance records are separate from login identities. A person can have no KEEN account and still have a role, assets, training or other assurance records. A vendor can supply assets; assigning an asset to a vendor moves its single supplier relationship. Do not turn an assurance record into an implicit authentication grant.

ISMS objectives, measures, metrics, assets, access records and meetings link assurance activity to the framework. The access matrix records intended or assessed access; it does not implement access in external systems. Scheduled audits and metric tasks run through the worker and need deterministic, idempotent behaviour.

9. Frontend conventions and documentation

Pages are static HTML with an ES-module controller. Use the shared API helpers for credentials, errors and CSRF headers. Escape plain values with esc, validate external links with safeExternalHref, and render rich text through the shared sanitiser. D3 labels should use text content rather than untrusted markup.

Carry the active framework through relevant queries and links. Avoid treating global objects, such as the underlying risk scenario, as though switching framework creates another object. Keep accessible labels, status messages and keyboard interaction when adding editors.

The documentation package generates the three online manuals from the same Markdown sources as the Word manuals. Static HTML, CSS, a small local search script and local screenshots require no external analytics or CDN. The KEEN UI exposes a Documentation link and serves the pages under /help/. The hosting portal uses the same assets under /help/.

To update a workflow, edit its canonical source, regenerate both formats, inspect the Word pages and check online anchors/search. Screenshots illustrate a demo state; instructions should still name the controls and describe the intended result. Do not publish customer-specific screenshots or credentials.

10. Develop and verify a change

  1. Identify the authoritative model, route, permission, UI controller and migration for the feature.
  2. Add schema changes through Alembic and update models together. Preserve existing data and relationships.
  3. Validate payloads and enforce permissions at the route and middleware boundaries.
  4. Keep multi-record updates transactional and use version checks where another editor can race.
  5. Record semantic changes, invalidate derived data and bump authorisation versions when relevant.
  6. Update the corresponding UI and manual from the actual behaviour.
  7. Test the consequential boundaries: unauthorised access, concurrent edits, invalid relationships, replay, rollback and historical data handling.

Use PostgreSQL integration tests for database-specific behaviour, including JSONB, row locking, uniqueness and transaction semantics. SQLite tests can cover isolated helpers but cannot prove PostgreSQL concurrency. For hosted changes, include negative cross-tenant identity and storage checks using separate test tenants.

Do not use a production demo as a write-test environment. Rehearse migrations and restores with disposable data, pin deployment dependencies, and record the source/image versions used in a release. Full UI image builds require the matching design-system dependency.

Documentation edition: 2 October 2026