CLI

falllow is how work is submitted and followed. Progress goes to stderr and the command's output to stdout, so it pipes the same as the local command. Every command takes --json.

Commands

CommandWhat it does
loginStore the server and a token for this machine, verifying them immediately.
runSubmit a job and stream it. Everything after -- is the command.
psRecent runs with status, exit code and machine.
logsPrint a run's output, or -f to follow it.
waitBlock until a run settles, then exit with its code.
showOne run in detail, including why it did not complete.
cancelStop runs. --abort kills instead of stopping gracefully.
devicesThe machines in the pool and how many of their blocks are busy.

Running things

falllow run
# stream it to completion, exit with its code
falllow run -- pytest -q

# ask for more of a machine
falllow run --cpu 8 --memory 16G -- cargo test --release

# a GPU and an image with the toolkit in it
falllow run --gpu 24GB --image nvidia/cuda:12.4.1-devel-ubuntu22.04 -- ./train.sh

# no sources needed
falllow run --no-workspace -- nproc

# hand back the run name and carry on
falllow run --detach -- ./long-solve.sh
FlagMeaning
--cpuCores the job needs at least. The job sees it as FALLLOW_CORES. Default 1.
--memoryMemory it needs at least, e.g. 8G. Default 1G.
--diskDisk it needs at least. Default 10G.
--gpuA GPU, by memory or name: 24GB, A100:2.
--imageContainer image to run in. Without it, dstack's base image.
--env, -eKEY=VALUE, or a bare KEY to pass on its value here. Repeatable.
--timeoutWall clock limit once running, e.g. 30m or 2h. Default 7d.
--queue-timeoutHow long to wait for a free machine. Default 1d.
--priority0 to 100, higher is placed first.
--shardSplit into this many independent jobs.
--nameRun name. Generated from the command if left out.
--notifyPOST the outcome as JSON to this URL when the run settles, from the server. Also on wait, from there.
--detach, -dPrint the run name instead of streaming to completion.
--includeShip this path even when an ignore rule excludes it. Repeatable.
--no-vcs-ignoreShip the directory as it lies, without git's ignore rules.
--no-workspaceShip no sources, for commands that need none.

The working directory

What is on disk is what runs. The enclosing git repository, or the directory itself if there is none, is packed on your machine and unpacked on the pool's, committed or not. Nothing is cloned from anywhere, so the pool needs no way into GitHub and no token travels with the job.

.git always stays home, so git describe does not work inside a job, and so does everything .dstackignore names. .gitignore decides inside a git repository, where it means build output, caches and virtualenvs. Outside a repository it does not: a directory assembled by hand is meant to travel as it lies, and a .gitignore copied in with the sources would otherwise keep excluding files where there is no git at all.

What a rule leaves out is named when the job is sent, not discovered as a missing file in a container. --include PATH ships a path however it was ignored, repeatable, and --no-vcs-ignore ships the whole directory without git's rules: for the wheel built for the pool, the model, the data package. The server accepts 256 MB per upload; a tree bigger than that wants a .dstackignore, or the large part can stay on the machine instead of travelling with every job.

The directory lands at /workspace, and the command runs from the subdirectory you typed it in: falllow run -- make in crates/solver runs in /workspace/crates/solver.

Sharding

--shard submits several independent jobs rather than splitting the work itself, only your command knows how to divide it. Each shard is told its slice and decides what that means.

Each shard is a run of its own, so --cpu and --memory are what one shard gets, not what they share: eight shards at --memory 4G ask the pool for 32 GB in total. They do not have to be on the same machine, or run at the same time; whatever does not fit waits its turn. A shard that exceeds its memory is killed on its own, and the others carry on.

Eight ways at once
falllow run --shard 8 -- pytest -q
What the command does with it
import os

index = int(os.environ.get("FALLLOW_SHARD_INDEX", 0))
count = int(os.environ.get("FALLLOW_SHARD_COUNT", 1))

mine = [t for i, t in enumerate(tests) if i % count == index]

The shards queue like any other jobs and run wherever there is room. Their output is printed one shard after another as they finish, since interleaved lines from eight jobs read as none. The exit code is the first genuine failure among them; an infrastructure failure only surfaces when no shard actually failed.

Waiting for a machine

A job that finds every fitting machine busy waits rather than failing. dstack looks at a waiting run again after a delay that grows: 15s, 30s, 1m, 2m, 5m, then every 10 minutes. The last step is worth knowing, because a machine that has just been freed can sit idle for minutes with something in the queue. falllow queue shows what is waiting, in the order the server considers it, and when each one is due to be tried again.

falllow queue
$ falllow queue
 #  RUN            PRIO  WAITING TRIES  NEXT TRY  NEEDS
 1  sweep-a          90       4m     5       12s  16 cores
 2  pytest-77c2       0      12m     6        4m   2 cores
falllow: place one in front with: falllow bump <run>

Order is priority first, then whoever has gone longest without being looked at. --priority 90 on run asks to be placed before the rest, and falllow bump <run> does the same for something already waiting: it queues the same job again at a higher priority and stops the copy that was waiting. Nothing is uploaded twice, the archive already on the server is reused. Priority decides who is considered first rather than who wins: placement happens in parallel, so two runs that become eligible together race for the machine.

What a job leaves behind

Every run gets a directory of its own on the machine it runs on, mounted at /artifacts and named in FALLLOW_ARTIFACTS. It outlives the container, so what a job wrote is still there when the run is over, including when the run failed. Nothing is collected by pattern: what the job put there is what it kept.

falllow fetch <run> brings it back, together with the run's log, into artifacts/<run>/. It reads through the pool's server, so it works from a laptop or from CI without being on the pool's network; on the network the directory can be read straight off the machine.

Each run writes into a directory of its own, its name plus a token, and the server files what it collects under the run's id. Names repeat: --name nightly is the same name every night, and dstack frees a name once its run has ended, so a directory keyed by the name alone would let one night's results overwrite the last. falllow fetch <run> takes a name, meaning the most recent run that had it, or an id from falllow show --json, meaning one particular run for ever.

Once a run settles the server collects its artifacts onto its own disk. A machine is not a safe place for the only copy: a rented one is switched off between jobs and a sponsored one is returned when the grant ends, and neither is a reason to lose what a run produced. So fetching keeps working after the machine is gone. A run too large to collect stays where it was made and says so; the collection as a whole stays inside a budget, oldest dropped first, and falllow devices prints what it takes up. falllow rm <run> deletes both copies.

falllow fetch
$ falllow fetch sweep-sh-9c1e
falllow: 4 files in artifacts/sweep-sh-9c1e

$ ls artifacts/sweep-sh-9c1e
falllow.log   result.json   plots/sweep.png   plots/error.png

The environment a job sees

VariableValue
FALLLOW_CORESThe cores asked for. Pass it on: make -j$FALLLOW_CORES.
FALLLOW_MEMORY_MIBThe memory asked for, in MiB. free reports the whole machine, so size by this.
FALLLOW_ARTIFACTSWhere to put what should outlive the job. Fetched with falllow fetch.
MAKEFLAGS, OMP_NUM_THREADS, and friendsSet to the same number, so tools that read them keep to the share the job was given. A job gets a CPU quota, not a smaller machine, so nproc still answers with the whole machine. Anything passed with -e wins.
FALLLOW_SHARD_INDEXThis shard's index, zero based. Only when sharding.
FALLLOW_SHARD_COUNTHow many shards there are.

For scripts and agents

The contract is the exit code and --json; nothing needs parsing from the human output. The simplest use is the blocking one: falllow run in place of the command, then branch on the exit code.

Branching on the outcome
falllow run --cpu 4 -- pytest -q
case $? in
  0)   echo "green" ;;
  170) echo "the pool failed, not the tests, safe to run again" ;;
  130) echo "somebody stopped it" ;;
  *)   echo "the tests failed" ;;
esac

For anything long, start it detached and come back. The run name is the handle; every command takes it, and --json gives the same object everywhere.

Start, check, wait
name=$(falllow run -d -- ./long-solve.sh)   # prints the run name at once
falllow show "$name" --json                 # status, exit_code (null until settled), machine
falllow wait "$name" --json                 # blocks, prints the outcome, exits with its code
falllow logs "$name" -f                     # the output, following
falllow wait --json
$ falllow wait --json cargo-test-3f9a
{
  "name": "cargo-test-3f9a",
  "status": "failed",
  "exit_code": 1,
  "exit_status": 1,
  "termination_reason": "container_exited_with_error",
  "submitted_at": "2026-09-16T09:38:50.980859+00:00",
  "finished_at": "2026-09-16T09:39:32.780373+00:00",
  "instance": "10.7.0.3",
  "attempts": 1
}

exit_code is what falllow run would have exited with, null while the run is unsettled. exit_status is the command's raw status if it got to run. termination_reason says why a run ended when it was not the command's doing.

To be told rather than to poll, --notify on run posts the outcome to a URL when the run settles. The URL is carried on the run and the pool's server does the posting, so a job started with -d is announced long after the terminal that started it is closed. The body is the same JSON --json prints, plus a one line text for webhooks that show a message and nothing else. On wait the flag posts from the waiting process instead, since a run already under way cannot be told where to report.

Getting told
falllow run -d --notify https://hooks.example.com/falllow -- ./sweep.sh
# the server POSTs the same JSON that --json prints, once the run settles,
# whether or not anything is still attached
falllow ps
$ falllow ps
NAME                             STATUS         EXIT  INSTANCE           SUBMITTED
sweep-sh-9c1e-1                  running              10.7.0.2           2026-09-16T09:41:55
cargo-test-3f9a                  failed            1  10.7.0.3           2026-09-16T09:38:50
pytest-77c2                      done              0  10.7.0.3           2026-09-16T09:38:26

Exit codes

CodeMeaning
0The command succeeded.
anything elseThe command's own exit code, passed through unchanged.
130Somebody cancelled the run.
137The kernel killed the container, nearly always because it used more memory than it asked for.
170Reserved: the pool failed, not your command. No machine free in time, a machine lost, the timeout passed.

Ctrl-C while following detaches and leaves the job running; falllow logs -f picks it up again.