Documentation

falllow turns the machines you own and the boxes you rent into one compute pool. You submit a command, it runs on whichever machine has room, and it behaves as if it had run locally.

Quickstart

You need a server and at least one machine in the fleet, Setup covers both. With those in place, submitting work is one command.

Your first job
# once, on your machine
falllow login --server https://app.falllow.com --token <token>

# from any git working directory
falllow run -- pytest -q

The pieces

  • The server, dstack with Postgres behind Caddy. It holds the queue, decides placement, keeps the logs, and serves the API the CLI and the manager talk to.
  • The fleet, the machines jobs run on. Each is a Linux host with Docker that the server reaches over SSH. The server installs a small agent on it and runs every job in a container. The machines reach nothing of yours: not GitHub, not each other.
  • The CLI (falllow), how work is submitted and followed. A thin layer over dstack's client, built to be driven by scripts and agents as much as by hand: every command takes --json.
  • The manager, the web app at the server's address. Jobs with their live output and attempts, and the machines with what they are running.

What a job is

A command, your working directory and a resource request. The directory is packed on your machine and unpacked on the pool's, so the job sees it exactly as you have it, committed or not. The command runs from the same subdirectory you typed it in, and its output streams back while it happens. A job never spans machines.

A machine is divided into blocks, and a job takes the blocks its request needs. A machine with four blocks runs up to four jobs side by side. When every machine that fits is busy, the job waits in the queue until one is free.

States

StateMeaning
queuedAccepted, waiting for a machine with a free block.
preparingPlaced. The container image is pulled and the tree applied.
runningThe command is executing.
succeededExited zero.
failedExited non-zero.
infra failureNever ran or was cut short: no machine within the queue timeout, a machine lost.
cancelledStopped on request.

Failure means two different things

A test suite that failed and a machine that vanished are both "the job did not succeed", and treating them the same is how you either run a real failure again until it happens to pass, or give up on work that only lost its machine.

OutcomeCauseExit code
failedThe command exited non-zero. Its verdict.The command's own
infra failureNo machine free within --queue-timeout, the machine became unreachable, --timeout passed.170
cancelledSomeone stopped it.130

So falllow run -- pytest substitutes for pytest in a script: the exit code you get is the one you would have got locally. 170 is reserved and means only that the pool failed you.

Where work goes

Only to machines already in the fleet. The server can also start cloud instances on demand, but falllow does not let a busy pool decide to spend money: a job waits for a machine of yours instead. Renting more is a change to the fleet, made on purpose.

A job that no machine in the pool could ever take, more memory than any machine has, a GPU where there is none, is refused when you submit it rather than queued.

Keep reading

  • Setup, the server, the network and the machines.
  • CLI, every command, the environment a job sees, and the JSON output.
  • CI, running a workflow's heavy steps on the pool.