PoolsA pool is a named set of files belonging to a repo, that the builders running its tasks keep synced through sai-server. It's meant for state that work builds up over many tasks and many builders, like fuzzing corpora: every builder working on the repo's fuzzing adds to one shared corpus, and starts from what all the others found. Asking for a pool in .sai.jsonA configuration names the pool its tasks use: The name is up to 32 lowercase letters, digits, What the task seesThe builder keeps its copy of the pool under
The builder syncs the pool before the task's first build step starts (unless it did within the last minute), every minute while the task runs, and once more after it ends. The builder does it rather than the task, so a task that's stopped, eg, because real work needed the builder, loses nothing it wrote. Several tasks on the builder using the same pool share one copy. Files in SAI_POOL_DIROnly files at
That's how libFuzzer names its corpus files already, so a libFuzzer corpus
dir per target, eg, Files arriving from the server are written somewhere else first and renamed into place, so a fuzzer reading the directory never sees one half written. FindingsFiles at The server groups findings into bugs, and publishes each bug's reproducer in
Replacing a sub, eg, after minimizing a corpusTasks can't delete shared files, since deleting one locally doesn't mean the
others should lose it. To replace everything in a sub instead, eg, with a
corpus libFuzzer's
At the next sync, the server removes every file in the sub it had up to that number that isn't in the directory now, and the other builders remove them when they next sync. Files anyone added after that number are kept, since the task didn't know about them. SizesA synced file can be up to 1MiB, and a pool holds up to 2 million files. On the serverEach repo's pool is a sqlite db next to the others,
TrustOnly a builder that has the fleet link key can sync, and only the pool of a task it was given: it has to present the task's upload nonce. But what goes in a pool comes from the tasks, which run what's pushed to the repo, so a pool is only as trustworthy as the people who can push to it. Content addressed files are checked against their names on both ends, names are checked before they're used as paths, and sizes are capped. |