CLI
falllow is how work is submitted and followed. Progress goes to stderr and the
command's output to stdout, so it pipes the same as the local command. Every command takes --json.
Commands
| Command | What it does |
|---|---|
login | Store the server and a token for this machine, verifying them immediately. |
run | Submit a job and stream it. Everything after -- is the command. |
ps | Recent runs with status, exit code and machine. |
logs | Print a run's output, or -f to follow it. |
wait | Block until a run settles, then exit with its code. |
show | One run in detail, including why it did not complete. |
cancel | Stop runs. --abort kills instead of stopping gracefully. |
devices | The machines in the pool and how many of their blocks are busy. |
Running things
# stream it to completion, exit with its code
falllow run -- pytest -q
# ask for more of a machine
falllow run --cpu 8 --memory 16G -- cargo test --release
# a GPU and an image with the toolkit in it
falllow run --gpu 24GB --image nvidia/cuda:12.4.1-devel-ubuntu22.04 -- ./train.sh
# no sources needed
falllow run --no-workspace -- nproc
# hand back the run name and carry on
falllow run --detach -- ./long-solve.sh| Flag | Meaning |
|---|---|
--cpu | Cores the job needs at least. The job sees it as FALLLOW_CORES. Default 1. |
--memory | Memory it needs at least, e.g. 8G. Default 1G. |
--disk | Disk it needs at least. Default 10G. |
--gpu | A GPU, by memory or name: 24GB, A100:2. |
--image | Container image to run in. Without it, dstack's base image. |
--env, -e | KEY=VALUE, or a bare KEY to pass on its value here. Repeatable. |
--timeout | Wall clock limit once running, e.g. 30m or 2h. Default 7d. |
--queue-timeout | How long to wait for a free machine. Default 1d. |
--priority | 0 to 100, higher is placed first. |
--shard | Split into this many independent jobs. |
--name | Run name. Generated from the command if left out. |
--notify | POST the outcome as JSON to this URL when the run settles, from the server. Also on wait, from there. |
--detach, -d | Print the run name instead of streaming to completion. |
--include | Ship this path even when an ignore rule excludes it. Repeatable. |
--no-vcs-ignore | Ship the directory as it lies, without git's ignore rules. |
--no-workspace | Ship no sources, for commands that need none. |
The working directory
What is on disk is what runs. The enclosing git repository, or the directory itself if there is none, is packed on your machine and unpacked on the pool's, committed or not. Nothing is cloned from anywhere, so the pool needs no way into GitHub and no token travels with the job.
.git always stays home, so git describe does not work inside a job,
and so does everything .dstackignore names. .gitignore decides inside
a git repository, where it means build output, caches and virtualenvs. Outside a repository it
does not: a directory assembled by hand is meant to travel as it lies, and a .gitignore copied in with the sources would otherwise keep excluding files where
there is no git at all.
What a rule leaves out is named when the job is sent, not discovered as a missing file in a
container. --include PATH ships a path however it was ignored, repeatable, and --no-vcs-ignore ships the whole directory without git's rules: for the wheel built
for the pool, the model, the data package. The server accepts 256 MB per upload; a tree bigger
than that wants a .dstackignore, or the large part can stay on the machine
instead of travelling with every job.
The directory lands at /workspace, and the command runs from the subdirectory you
typed it in: falllow run -- make in crates/solver runs in /workspace/crates/solver.
Sharding
--shard submits several independent jobs rather than splitting the work itself, only
your command knows how to divide it. Each shard is told its slice and decides what that means.
Each shard is a run of its own, so --cpu and --memory are what one
shard gets, not what they share: eight shards at --memory 4G ask the pool for
32 GB in total. They do not have to be on the same machine, or run at the same time; whatever
does not fit waits its turn. A shard that exceeds its memory is killed on its own, and the
others carry on.
falllow run --shard 8 -- pytest -qimport os
index = int(os.environ.get("FALLLOW_SHARD_INDEX", 0))
count = int(os.environ.get("FALLLOW_SHARD_COUNT", 1))
mine = [t for i, t in enumerate(tests) if i % count == index]The shards queue like any other jobs and run wherever there is room. Their output is printed one shard after another as they finish, since interleaved lines from eight jobs read as none. The exit code is the first genuine failure among them; an infrastructure failure only surfaces when no shard actually failed.
Waiting for a machine
A job that finds every fitting machine busy waits rather than failing. dstack looks at a
waiting run again after a delay that grows: 15s, 30s, 1m, 2m, 5m, then every 10 minutes. The
last step is worth knowing, because a machine that has just been freed can sit idle for
minutes with something in the queue. falllow queue shows what is waiting, in the
order the server considers it, and when each one is due to be tried again.
$ falllow queue
# RUN PRIO WAITING TRIES NEXT TRY NEEDS
1 sweep-a 90 4m 5 12s 16 cores
2 pytest-77c2 0 12m 6 4m 2 cores
falllow: place one in front with: falllow bump <run>Order is priority first, then whoever has gone longest without being looked at. --priority 90 on run asks to be placed before the rest, and falllow bump <run> does the same for something already waiting: it queues
the same job again at a higher priority and stops the copy that was waiting. Nothing is
uploaded twice, the archive already on the server is reused. Priority decides who is
considered first rather than who wins: placement happens in parallel, so two runs that become
eligible together race for the machine.
What a job leaves behind
Every run gets a directory of its own on the machine it runs on, mounted at /artifacts and named in FALLLOW_ARTIFACTS. It outlives the container,
so what a job wrote is still there when the run is over, including when the run failed. Nothing
is collected by pattern: what the job put there is what it kept.
falllow fetch <run> brings it back, together with the run's log, into artifacts/<run>/. It reads through the pool's server, so it works from a
laptop or from CI without being on the pool's network; on the network the directory can be read
straight off the machine.
Each run writes into a directory of its own, its name plus a token, and the server files
what it collects under the run's id. Names repeat: --name nightly is the same
name every night, and dstack frees a name once its run has ended, so a directory keyed by
the name alone would let one night's results overwrite the last. falllow fetch <run> takes a name, meaning the most recent run that had it,
or an id from falllow show --json, meaning one particular run for ever.
Once a run settles the server collects its artifacts onto its own disk. A machine is not a safe
place for the only copy: a rented one is switched off between jobs and a sponsored one is
returned when the grant ends, and neither is a reason to lose what a run produced. So fetching
keeps working after the machine is gone. A run too large to collect stays where it was made and
says so; the collection as a whole stays inside a budget, oldest dropped first, and falllow devices prints what it takes up. falllow rm <run> deletes both copies.
$ falllow fetch sweep-sh-9c1e
falllow: 4 files in artifacts/sweep-sh-9c1e
$ ls artifacts/sweep-sh-9c1e
falllow.log result.json plots/sweep.png plots/error.pngThe environment a job sees
| Variable | Value |
|---|---|
FALLLOW_CORES | The cores asked for. Pass it on: make -j$FALLLOW_CORES. |
FALLLOW_MEMORY_MIB | The memory asked for, in MiB. free reports the whole machine, so size by this. |
FALLLOW_ARTIFACTS | Where to put what should outlive the job. Fetched with falllow fetch. |
MAKEFLAGS, OMP_NUM_THREADS, and friends | Set to the same number, so tools that read them keep to the share the job was
given. A job gets a CPU quota, not a smaller machine, so nproc still
answers with the whole machine. Anything passed with -e wins. |
FALLLOW_SHARD_INDEX | This shard's index, zero based. Only when sharding. |
FALLLOW_SHARD_COUNT | How many shards there are. |
For scripts and agents
The contract is the exit code and --json; nothing needs parsing from the human
output. The simplest use is the blocking one: falllow run in place of the command,
then branch on the exit code.
falllow run --cpu 4 -- pytest -q
case $? in
0) echo "green" ;;
170) echo "the pool failed, not the tests, safe to run again" ;;
130) echo "somebody stopped it" ;;
*) echo "the tests failed" ;;
esacFor anything long, start it detached and come back. The run name is the handle; every command
takes it, and --json gives the same object everywhere.
name=$(falllow run -d -- ./long-solve.sh) # prints the run name at once
falllow show "$name" --json # status, exit_code (null until settled), machine
falllow wait "$name" --json # blocks, prints the outcome, exits with its code
falllow logs "$name" -f # the output, following$ falllow wait --json cargo-test-3f9a
{
"name": "cargo-test-3f9a",
"status": "failed",
"exit_code": 1,
"exit_status": 1,
"termination_reason": "container_exited_with_error",
"submitted_at": "2026-09-16T09:38:50.980859+00:00",
"finished_at": "2026-09-16T09:39:32.780373+00:00",
"instance": "10.7.0.3",
"attempts": 1
}exit_code is what falllow run would have exited with, null while the
run is unsettled. exit_status is the command's raw status if it got to run. termination_reason says why a run ended when it was not the command's doing.
To be told rather than to poll, --notify on run posts the outcome to a
URL when the run settles. The URL is carried on the run and the pool's server does the posting,
so a job started with -d is announced long after the terminal that started it is
closed. The body is the same JSON --json prints, plus a one line text for webhooks that show a message and nothing else. On wait the
flag posts from the waiting process instead, since a run already under way cannot be told where
to report.
falllow run -d --notify https://hooks.example.com/falllow -- ./sweep.sh
# the server POSTs the same JSON that --json prints, once the run settles,
# whether or not anything is still attached$ falllow ps
NAME STATUS EXIT INSTANCE SUBMITTED
sweep-sh-9c1e-1 running 10.7.0.2 2026-09-16T09:41:55
cargo-test-3f9a failed 1 10.7.0.3 2026-09-16T09:38:50
pytest-77c2 done 0 10.7.0.3 2026-09-16T09:38:26Exit codes
| Code | Meaning |
|---|---|
0 | The command succeeded. |
| anything else | The command's own exit code, passed through unchanged. |
130 | Somebody cancelled the run. |
137 | The kernel killed the container, nearly always because it used more memory than it asked for. |
170 | Reserved: the pool failed, not your command. No machine free in time, a machine lost, the timeout passed. |
Ctrl-C while following detaches and leaves the job running; falllow logs -f picks it
up again.