When async Rust starts blocking

Updated

You can get pretty far building a Rust API without understanding the async runtime. Pick Axum, connect SQLx and Redis, and write handlers that await their dependencies. Requests are fast, and you focus on the application.

Then the workload grows. Queries return more rows, cached documents get larger, and imports process thousands of records. Suddenly, expensive requests affect cheap ones too. That is when the runtime deserves a closer look.

When a fast cache slows down the service

Consider an async handler that reads compressed JSON from Redis. Using illustrative helpers:

let compressed = cache.get("catalog").await?;
let json = decompress(&compressed)?;
let catalog = serde_json::from_slice::<Catalog>(&json)?;

For a small document, this works well. As it grows, decompression and parsing become expensive. Redis is still fast, but requests elsewhere in the service sometimes slow down too.

Those last two lines run synchronously on a Tokio worker. Here, blocking means keeping that worker occupied so it cannot run another task—whether it is computing or waiting synchronously. An async handler can do both.

This is one possible explanation for simultaneous slow requests. Confirm that they overlap on the same instance; similar latency alone does not establish the cause.

How Tokio runs your code

Calling an async function produces a future, which holds the state of the computation. Tokio asks a task to make progress by polling it. The task either finishes or returns control until it can continue. These outcomes are called Ready and Pending; a wakeup tells Tokio when to poll again.

Tokio runs tasks on worker threads. While a poll executes, it occupies its worker. When the task suspends waiting for I/O, that worker can run something else.

An .await can suspend execution, but it does not have to. If the awaited operation is already ready, execution continues. Tokio’s cooperative scheduling cannot interrupt arbitrary synchronous code halfway through a parser or loop.

In the cache example, the network wait can suspend. Decompression and JSON parsing cannot. If they take 300 milliseconds, that worker is occupied for 300 milliseconds. Other workers may continue, but fewer are available to serve requests.

Tokio provides different ways to organize that work:

Operation What happens
operation.await Advances the operation and suspends if it returns Pending; ready work continues immediately.
tokio::spawn(...) Schedules another async task on the runtime’s worker pool.
yield_now().await Gives the scheduler an opportunity to run other work before this task continues.
spawn_blocking(...) Runs a synchronous closure on a separate blocking pool; awaiting its handle lets the caller suspend until it finishes.

Yielding breaks a loop into shorter turns on the async runtime. Offloading moves synchronous work to another pool. Neither makes the computation cheaper, and tokio::spawn alone does not isolate blocking code.

What can occupy a worker

  • Parsing and encoding: JSON or CSV parsing, compression, image processing, and response serialization.
  • Database results: draining buffered rows, decoding large values, and converting them into application objects after the network wait.
  • Bulk processing: sorting, deduplication, repeated collection scans, or downloading in batches but processing everything at once.
  • Allocation and cleanup: cloning large objects, building temporary structures, and freeing them when they are dropped.
  • Synchronous waits: blocking file or network APIs, std::thread::sleep, and contended synchronous locks. These can occupy a worker with little CPU usage.
  • Diagnostics: backtrace capture, symbol resolution, and synchronous log output on a request path.

These operations are not inherently problematic. Their duration, input size, frequency, and execution thread determine whether they interfere with other requests.

Choose the response based on the work:

  • Unnecessary work: remove it first. Fetch fewer fields, bound results, and avoid repeated scans or copies.
  • Blocking I/O: prefer an async API when available.
  • One expensive synchronous call: use spawn_blocking with bounded concurrency.
  • Many small operations: process bounded batches and yield between them.

Move expensive work off async workers

For substantial synchronous work that remains, use spawn_blocking. Extract the cache decompression and parsing into a synchronous helper:

let compressed = cache.get("catalog").await?;
let catalog = tokio::task::spawn_blocking(move || {
    decode_catalog(compressed)
})
.await??;

decode_catalog contains the same decompression and JSON parsing shown earlier and returns a Result. The two ? operators propagate join and decoding errors. The work now runs on the blocking pool while the handler awaits its result.

Offloading still consumes CPU. Limit concurrent expensive jobs, for example with a shared semaphore acquired asynchronously before spawning. Keep the permit inside the closure until it finishes. A started blocking task cannot be stopped by aborting its handle. Tokio’s spawn_blocking documentation explains these constraints.

Keep network and database waits async. Offload the substantial synchronous phase; tiny operations rarely justify a thread handoff.

Let long loops yield

Sometimes each operation is cheap, but a loop performs thousands without suspending. If you control the loop, yield between small batches:

for batch in records.chunks(128) {
    for record in batch {
        process(record);
    }
    tokio::task::yield_now().await;
}

Choose the batch size by measured processing time; 128 is only an example. yield_now gives other tasks a chance to run, but Tokio may immediately poll the same task again. It cannot interrupt an expensive process(record) call.

Alternatively, consume_budget yields when Tokio’s cooperative budget is exhausted. That budget counts participating operations, not milliseconds.

The same issue can arise with SQL streams: rows.try_next().await may repeatedly return ready when rows are buffered. Streaming alone does not ensure short polls, and yielding after fetch_all().await is too late to interrupt work inside it.

Check whether other work can still run

Request duration and poll duration answer different questions. A request can wait seconds for a database while yielding promptly. Another can finish sooner but occupy a worker for most of its execution.

Measure individual polls and run a periodic heartbeat to detect scheduling delays. These measure elapsed time, not necessarily CPU usage: blocking waits and operating-system scheduling also contribute.

Four tools help at different stages; their linked guides cover setup:

  1. Start with tokio-blocked to log warnings when task polls exceed a configured duration. It reports the task’s spawn location, which may not identify the blocking function inside it.
  2. Use tokio-console to inspect live tasks and their poll durations interactively.
  3. Use cargo-flamegraph to locate CPU-heavy functions. CPU samples alone will not explain time a thread spent blocked or descheduled.
  4. Use tokio-metrics to track poll times and scheduling delays over time for instrumented futures.

Both tokio-blocked and tokio-console require Tokio’s tracing instrumentation and a build with tokio_unstable. The drawback is compatibility: these unstable features can change even within Tokio 1.x releases, so check your tooling when upgrading. They require an instrumented build; you cannot simply attach them to an unmodified application.

Test large representative inputs in a release build, with a cheap request running concurrently. Check its latency alongside the expensive operation. A faster local CPU can hide problems that appear in production.

For our cache example, decoding may take just as long after offloading. The improvement to look for is less interference with unrelated requests.

You do not need to offload every calculation. Start with work that grows with input, and ask: how much can this task do before it gives control back?