docs/ss-policy.md

Secure Streams policy

Draft, for iteration. It records what was decided with Andy on 2026-10-07 and 2026-10-08, and marks what is still open. The stream and service API it configures is a separate document, still to be written (ss-design.md); io-model.md is the layer underneath.

What this replaces

C lws has two configuration languages that grew apart:

  • lwsws' lejp-conf (lib/io/lejp-conf.c). It composes a server: vhosts with listeners and tls, protocols enabled per vhost with their options (pvo), and mounts that place them in the vhost's URL space (file://, cgi://, proxy, redirect, callback://).
  • The Secure Streams policy (lib/secure-streams/policy-json.c). It describes streamtypes: endpoints, protocol, tls trust, retry, metadata. Its server side is much smaller: a streamtype with "server": true is one listener and one handler for every path, which is what lejp-conf would call a single protocol plugin.

npro has one: the SS policy, grown to do lejp-conf's job as well. SS is npro's user API for client and server, and its policy is the only configuration language.

Principles

  • The policy is Rust types. JSON (parsed by npro-json) is one way to build them; Rust code is another. A const policy in code needs no code generator, so C's policy2c and static-policy builds have no equivalent: an embedded build simply has no JSON.
  • Most programs need no policy at all. A stream made from a URL derives what C's built-in __default streamtype gives: protocol and tls from the scheme, the default port, the system's trust store, redirects followed, headers read and written by name. A policy carries deployment decisions (endpoints, trust, certs, retry, listeners, routes) for when someone other than the programmer should make them.
  • Parameters only. Configuration binds declared parameters; nothing else can be changed from outside.
  • Merging is by namespace, never by file order. There is no equivalent of conf.d's alphabetical loading.
  • Nothing is accepted and ignored. An unknown key, a parameter bound to the wrong type or out of its bounds, or a required parameter left unbound, fails resolution, and the error names its path, for example `vhosts."warmcat.com".routes[2].to: no endpoint "mirror.wss"`.
  • npro is not an orchestrator. It defines the policy, resolves it and serves it to the processes that need it. Starting, restarting and placing processes is systemd's job, or whatever the deployment uses.

Parameters

A parameter is a declared, typed setting: its name, its type, its default or that it is required, its bounds, and one line of documentation. Configuration does nothing but bind parameters.

Parameters are declared by:

  • npro itself, under the top-level npro namespace. Every configuration struct in the core is declared this way: h1 header limits, the ws rx buffer and the permessage-deflate message cap, keepalive and other timeouts, tls ciphers, retry defaults. Bounds come from C's ranges. Names are not settled; for illustration, npro.h1.header-bytes, npro.ws.pmd.max-message, npro.tls.ciphers.
  • each service, in its structure (below): its options, and the bind address of each of its named listeners, which is always a parameter.

The declarations are written once in Rust, in a macro_rules! table beside the configuration struct, with no serde and no proc-macros, and no_std is unaffected. Two things come from them:

  • the decoder from the policy into the struct, so a declared parameter can never be silently ignored;
  • the reference documentation of every parameter, generated, so it cannot drift from the code as lejp-conf's README has.

Scopes

A parameter is bound at a scope: the whole process, a listener, a vhost, a route, or an instance. The rules:

  • binding the same parameter twice in one scope is an error;
  • a nested scope refines its parent for everything inside it: a route inside a vhost, a vhost inside a listener, an instance inside its deployment.

Precedence is decided by the structure, never by which file was read last. (Open: agreed in discussion, not yet confirmed.)

Exports and references

Some facts belong to one layer and are needed by another: the front's public host name, which a webrtc instance announces in its URLs; its external address, which ICE needs. The owner exports them by name, and a binding elsewhere refers to them:

"public-host": "${front.public-host}"

The dependency is then visible where it is used, and no layer reaches into another's namespace. An unknown reference fails resolution. This is the only substitution there is: lejp-conf's "=NAME" / ${NAME} preprocessor is not carried over.

The layers

Taking a webrtc service as the example, which runs as its own process and has both ws and udp listeners:

layerwritten bychanges whensays
structurethe service's developer; it ships with the cratethe code changesits named listeners and what each is ("signalling": ws, "media": datagram), with no addresses; its parameters, with their types, defaults and bounds; the client streamtypes it uses
deploymentwhoever runs itinstances or machines changethe instances (webrtc.eu1, webrtc.eu2), and for each, the bind address of each named listener and its parameter values
frontwhoever runs the public facepublic names or certs changelisteners, tls, vhosts, routes, guards, the exports, and npro.* parameters at any scope
localwhoever runs the machinethe machine changesthe process's identity, where its authority is, the key that checks what the authority sends, and the process's own secrets

The deployment and front layers are read by the front; how the operator splits them into files is up to the operator, since they merge by namespace. The structure comes from the service: from its crate, for a service linked into the front, or from the instance itself, for one in another process (see "The handshake"). The local layer stays on its machine and never travels.

The instances are the namespace: webrtc.eu1.media.bind, webrtc.eu1.max-rooms. A shorthand for "n instances, ports from a base" may come later; instances are listed explicitly for now.

An example

The structure, as npro-webrtc declares it (shown as JSON, which is also how an instance sends it):

{
  "component": "npro-webrtc", "version": "0.3.0",
  "listeners": {
    "signalling": { "kind": "ws", "subprotocol": "lws-webrtc" },
    "media":      { "kind": "datagram" }
  },
  "params": {
    "max-rooms":   { "type": "u32", "default": 64, "min": 1, "max": 4096 },
    "public-host": { "type": "host", "required": true }
  }
}

The deployment:

{
  "instances": {
    "webrtc.eu1": {
      "component": "npro-webrtc",
      "bind":   { "signalling": "unix:/run/npro/webrtc-eu1.sock",
                  "media": "[::]:3478" },
      "params": { "max-rooms": 128, "public-host": "${front.public-host}" }
    }
  }
}

The front:

{
  "exports": { "public-host": "rtc.example.com" },
  "listeners": {
    "https": { "bind": "[::]:443", "tls": { "cert": "rtc" },
               "vhosts": ["rtc.example.com"] }
  },
  "vhosts": {
    "rtc.example.com": {
      "params": { "npro.h1.keepalive": "5s" },
      "routes": [
        { "path": "/",    "to": "www.http" },
        { "path": "/rtc", "to": "webrtc.eu1.signalling" }
      ]
    }
  },
  "services": {
    "www": { "type": "files", "params": { "root": "/var/www/rtc" } }
  }
}

The route names webrtc.eu1.signalling, not an address: resolution finds where that is from the deployment, so each address is written once. media faces the public directly; a datagram listener cannot be proxied by the front.

Resolution and publication

The front is the policy authority. At its start, and when it is told to reload, it resolves the layers. It then serves each instance its fragment: everything bound in that instance's namespace, flattened, with the exports it refers to and the npro.* parameters that apply to it.

An instance fetches its fragment when it starts, as an SS client stream to its authority. A restarted instance, on any machine, always starts with the latest configuration; there is no resolving at deploy time, which could only produce configuration that had gone stale.

  • The instance caches the last fragment it accepted. If the authority cannot be reached, it starts from the cache. With no cache it fails to start, and systemd tries again.
  • A new configuration reaches an instance when it next restarts. *(Later: the front could tell a connected instance that its namespace has changed, and the instance would exit cleanly, so systemd brings it back with the new one.)*

The handshake

The front links none of the separately built services, so it cannot know their structure by itself. **The instance sends its structure when it asks for its fragment**: "I am webrtc.eu1, built from npro-webrtc 0.3.0, and these are my listeners and parameters." The front resolves against that, and answers with the fragment or with the resolution's errors. So:

  • the front stays generic, and links only the services it hosts itself;
  • an instance whose new binary declares a new parameter gets its default, or a clear refusal if it is required and unbound;
  • the front knows which configuration version each instance took, which, with the health it sees by proxying, gives a status page such as "eu1: healthy, policy v42" or "eu2: refused, media.bind unbound".

(Open: agreed in discussion as the direction, not yet confirmed.)

Signing

A fragment is configuration, and configuration is authority: it can redirect endpoints and name certs. So:

  • the front signs each fragment, and the instance checks it with the key in its local layer. The front has to hold the signing key, since it resolves at run time. That adds no exposure: the front already holds the tls keys and sees all the traffic;
  • a fragment carries its namespace and a version, and an instance refuses one for another namespace, or older than the one it has cached, so a captured fragment cannot be replayed to roll an instance back;
  • fragments are signed, not encrypted, so **they carry no private keys**. An instance's own secrets (the private key of its mutual tls with the front, say) stay in its local layer, and a fragment refers to them by name.

The format (JWS, when npro has jose; or something smaller first) and the signature algorithm (Ed25519, from whatever crypto the tls decision admits) are open.

Transport to the authority

On the same machine, a unix domain socket whose permissions admit only the instance's user. Across machines, tcp with mutual tls, whose certs are named in the local layer. The same channel, in the other direction, is how the front proxies to the instance.

Health

systemd keeps processes running. The front sees whether each instance is healthy by proxying to it: failed connects and timeouts. It uses that for a fast 503 on the routes of an instance that is down, for backing off its reconnects with the instance's retry policy, and for the status page. An optional active probe may come later.

From the C SS policy

Cnpro
release, product, schema-versionthe policy's header; the version that rollback protection checks is the fragment's, set by the front
retrynamed retry schemes, as now; their limits (C: at most 8 backoff steps) become declared bounds
certs (base64 DER inline)public certs may still be inline; private keys come only from the local layer or a file, never inline
trust_storeskept, referenced by name
s (client streamtypes)client streamtypes, declared in a structure by the service or program that uses them, with parameters the policy binds (endpoint, trust, retry)
"server": true streamtypeslisteners, vhosts, services and routes. C's one-listener server is the smallest case, and a short form of it stays
metadata, ${metadata}stream metadata, copied, never held by pointer, so a stream's state can cross a process boundary; headers by name are the default, which C has only with direct_proto_str
options[] (parsed, never used in C)service parameters
overlay (modifies existing streamtypes only)the layers
replacing the policy (destroys every stream)restart into the new fragment
fetch_policyfetching the fragment from the authority
static policy, policy2ca const policy in Rust
auth, metricslater; not designed here

From lejp-conf

Where each group of lejp-conf's keys lands. "Param" means a declared parameter bound at the scope named.

Globals

lejp-confnpro
uid, gid, username, groupname, rlimit-nofilenot npro: systemd's User=, Group=, LimitNOFILE=
count-threads, count-async-threadsthe runner's parameters
plugin-dirgone: services are compiled in, or run as their own process
init-sslgone
server-string, timeout-secs, http-header-datanpro.* params, process scope
ip-limit-ah, ip-limit-wsinpro.* limit params
default-alpnlistener tls
quic-*npro.quic.* params
reject-service-keywordsa route guard
cpd-bypassthe client side's captive portal detection, a param
ws-pingpong-secsgone (deprecated in C)

Per vhost

lejp-confnpro
namethe vhost's key
port, interface, unix-socket, unix-socket-perms, noipv6, ipv6only, fo-listen-queuelisteners, which are separate from vhosts; a listener lists the vhosts it serves
host-ssl-key, -cert, -ca, -cert-grace-secslistener tls, chosen per vhost by SNI; the private key from the local layer or a file
ciphers, tls13-ciphers, ecdh-curve, ssl-option-set/clear, alpn, client-cert-required, ignore-missing-certlistener tls params
sts, redirect-http, allow-non-tls, allow-http-on-https, strict-host-check, sni-fallbacklistener and vhost params
access-log, keepalive_timeout, h2-half-closed-long-poll, quic-mtu, quic-preferred-addressesnpro.* params, vhost scope
headers[]the vhost's response headers
error-document-404the vhost's error documents
listen-accept-role, -protocol, apply-, fallback-listen-accept, onlyrawa listener's kind, and its Raw endpoint
disable-no-protocol-ws-upgradesa route rule
enable-client-ssl, client-ssl-*client trust and certs, at the scope of the streams that use them
ws-protocols[] (pvo)services, with their parameters
dht[]a dht service

Per mount

lejp-confnpro
mountpoint, exact-match, append-paththe route's path and how it matches; the longest matching path wins
origin file://, gzip://the files service, with precompressed variants as its param
origin http://, https://the proxy service: a route bridged to a client stream
origin >http://, >https://a redirect route
origin callback://a service's endpoint, by name
origin cgi://, cgi-*open: a cgi service, or not carried over
default, extra-mimetypes, interpret, cache-*the files service's params
auth-mask, basic-auth, interceptor-pathroute guards
headers[], keepalive-timeout, no-ws-upgradesroute params
pmo[]service params, route scope

Not carried over: the "=NAME" / ${NAME} preprocessor (exports and references cover what it was used for), and conf.d order.

Decided

  • 2026-10-07: the SS policy is npro's one configuration language, and takes over lejp-conf's job. Services in their own processes are fronted over unix domain sockets, or mutual tls across machines.
  • 2026-10-08: parameters only; npro's own configuration is parameters too, under the npro namespace.
  • 2026-10-08: instances fetch their fragment from the front when they start, and cache the last good one; there is no resolving at deploy time. Fragments are signed, and checked with a key held locally.
  • 2026-10-08: npro does not supervise processes; systemd does. npro knows their health from proxying to them.

Open

  • The scope rule (same-scope duplicates are errors; nested scopes refine).
  • The handshake, with the instance sending its structure.
  • Names for the npro.* parameters, and how they are grouped.
  • The fragment's signature format and algorithm.
  • Telling a running instance that its namespace changed.
  • cgi.
  • The JSON shapes above, which are illustrations, not a schema.