Benchmarks

One release-built Rust client measures every server on the same machine and workloads, separating server and transport costs from language-specific client overhead.

Latest run: · AWS Graviton5, 32 cores, 61.6 GB

Canonical results pending

C# has been added to the release-pinned benchmark matrix. Its performance figures will appear after a full run on the canonical c9gd.8xlarge host; the published charts below remain the prior comparable snapshot.

Release inputs

Every run is bound to the full commit ID behind the release tag selected for that run. These historical measurements may predate the SDK versions in the capability matrix; Iroh has not been measured in this snapshot.

ImplementationReleaseCommitStatus
Pythonv0.43.01da388f7bf58Published release
TypeScriptv0.22.0423558fa50c7Published release
Gov0.26.18e7c49fd0724Published release
Rustv0.23.2eaba89b5d60ePublished release
Javav0.23.1a0309fe28073Published release
C#v0.7.08afe27f8c5abCanonical benchmark pending
C++v0.1.40fe27a2870edPublished release

Performance at a glance

Every panel is shown up front. Lower is better for latency; higher is better for call rate and payload bandwidth.

Single-call latency by language

Mean latency for a tiny unary call with one client. Color compares languages within each transport, so HTTP remains useful without flattening the faster local transports.

LanguageDirect transportsHTTP
Subprocess / stdioStdio + shared memoryUnix socketTCP loopbackHTTP identityHTTP zstd
Python54.1 μs56.2 μs66.4 μs78.7 μs103.8 μs134.7 μs
TypeScript79.4 μs—93.4 μs102.6 μs83.3 μs108.9 μs
Go17.3 μs17.8 μs25.9 μs42.2 μs57.5 μs87.8 μs
Rust13.1 μs14.3 μs21.6 μs39.3 μs46.0 μs57.6 μs
Java15.0 μs15.9 μs24.7 μs27.4 μs53.6 μs73.8 μs
C++13.7 μs15.7 μs23.5 μs38.2 μs49.1 μs68.3 μs
Within each column:FastestWithin 1.5×Within 2.5×Over 2.5×

Scaling relative to one client

One client shows request rate; 2–16 clients show throughput relative to that baseline. Green is close to linear scaling; red indicates saturation.

Multi-core implementations

These servers can dispatch independent connections across multiple CPU cores.

Go
Small unary calls
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket38K1.40×2.70×2.83×2.49×
TCP loopback23.7K1.45×3.13×3.45×3.44×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity17.5K1.78×2.85×3.79×3.18×
HTTP zstd11.3K1.85×3.75×7.49×12.08×
1 MiB round trips
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket2.1K1.69×2.90×4.44×5.50×
TCP loopback2K1.60×2.89×4.38×5.68×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity1.2K1.38×2.56×3.73×4.15×
HTTP zstd250.52.15×4.25×8.07×12.63×
Rust
Small unary calls
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket45.7K1.98×3.98×7.87×15.75×
TCP loopback24.8K1.89×3.65×6.99×12.01×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity21.7K1.90×3.70×7.02×12.03×
HTTP zstd17.6K1.90×3.70×7.26×12.22×
1 MiB round trips
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket5.9K1.89×3.43×6.17×9.65×
TCP loopback4.2K1.97×3.76×6.81×11.11×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity928.31.87×3.26×5.56×7.16×
HTTP zstd510.31.69×3.85×7.26×12.83×
Java
Small unary calls
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket40.1K1.89×3.91×7.67×8.70×
TCP loopback35.6K1.93×3.70×7.18×7.36×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity18.4K1.98×3.93×7.57×11.14×
HTTP zstd13.1K2.03×4.06×7.71×11.48×
1 MiB round trips
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket2.4K1.95×3.68×6.30×7.61×
TCP loopback2.2K1.99×3.76×6.78×8.93×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity863.21.96×3.47×5.22×5.60×
HTTP zstd594.42.02×3.94×7.04×9.86×
C++
Small unary calls
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket42.6K1.90×3.76×7.33×13.68×
TCP loopback26.4K1.93×3.76×7.27×11.19×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity20.2K1.98×3.93×7.70×11.81×
HTTP zstd14.5K1.98×3.95×7.79×12.28×
1 MiB round trips
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket5.8K1.91×3.53×6.30×9.47×
TCP loopback4.4K1.93×3.77×7.04×10.93×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity1.6K1.92×3.43×4.34×4.84×
HTTP zstd7242.00×3.93×6.83×10.12×

Single-process, one-core implementations

Python and TypeScript raw socket servers are bounded by one process core, so their client-count curves are shown separately. Python HTTP uses 16 Granian worker processes.

PythonRaw socket server: single process, one CPU core; HTTP: 16 Granian workers.
Small unary calls
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket15.2K0.67×0.48×0.48×0.48×
TCP loopback12.5K0.78×0.56×0.56×0.55×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity9.7K2.00×3.98×7.50×11.37×
HTTP zstd7.7K1.93×3.87×7.68×11.16×
1 MiB round trips
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket1.2K1.11×1.18×1.21×1.14×
TCP loopback1.2K1.17×1.25×1.24×1.10×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity559.61.95×3.34×5.07×7.08×
HTTP zstd414.51.88×3.92×6.69×10.60×
TypeScriptRaw socket server: single process, one CPU core.
Small unary calls
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket11K1.04×1.04×1.03×1.02×
TCP loopback9.5K1.14×1.18×1.17×1.15×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity12K1.55×1.59×1.61×1.57×
HTTP zstd9.2K1.59×1.65×1.64×1.64×
1 MiB round trips
Socket transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
Unix socket1K1.10×1.10×1.07×1.01×
TCP loopback1K1.07×1.06×1.07×1.04×
HTTP transports
Transport1 clientreq/s2 clients4 clients8 clients16 clients
HTTP identity6961.79×2.10×2.04×2.01×
HTTP zstd530.51.61×1.93×1.92×1.95×

Single-client large-payload bandwidth by transport

16 MiB echo throughput with one outstanding call, counting bytes in both directions. Client and server are not CPU-pinned, so these are single-client rather than single-core results. Color compares languages within each transport; higher is better.

LanguageDirect transportsHTTP
Subprocess / stdioStdio + shared memoryUnix socketTCP loopbackHTTP identityHTTP zstd
Python1.26 GiB/s1.54 GiB/s1.77 GiB/s1.74 GiB/s547 MiB/s832 MiB/s
TypeScript827 MiB/s—1.11 GiB/s1.13 GiB/s763 MiB/s1.14 GiB/s
Go1.68 GiB/s3.22 GiB/s2.28 GiB/s2.16 GiB/s831 MiB/s481 MiB/s
Rust1.74 GiB/s3.35 GiB/s2.82 GiB/s2.92 GiB/s680 MiB/s700 MiB/s
Java1.83 GiB/s2.79 GiB/s2.45 GiB/s2.43 GiB/s796 MiB/s1.14 GiB/s
C++1.74 GiB/s3.35 GiB/s3.04 GiB/s2.78 GiB/s637 MiB/s630 MiB/s
Within each column:FastestAt least 65%At least 35%Under 35%

Detailed results

Rust client → release server. Choose a workload, transport, and available client count to inspect percentiles, CPU, and quality.

Workload

Transport

Clients

Single-call workloads have one client; load workloads expose 1, 2, 4, 8, and 16.

ServerClientsMeanP50P95P99CallsPayloadHost CPUQuality

What is excluded

In-process pipes, WSGI test clients, debug builds, unpinned working trees, first-call startup time, and results from shared CI hosts are not used for cross-language claims.

Methodology

  • Release builds from isolated checkouts pinned to the commits above.
  • Python HTTP uses pinned Granian 2.8.1 with 16 worker processes and one blocking thread each.
  • Java runs on Amazon Corretto 25.0.4 so negotiated shared memory uses the JDK 22+ FFM implementation.
  • Latency uses a one-second warm-up and five one-second rounds; unstable cases receive two extension rounds.
  • Mean, P50, P95, and P99 use all steady-state wall-clock samples—not the fastest sample.
  • Scaling uses fresh servers at 1, 2, 4, 8, and 16 clients, with three two-second rounds and two extensions when noisy.
  • Scaling records call rate, latency, and whole-host CPU utilization for unary and 1 MiB echo workloads.
  • Binary echo throughput counts application payload in both request and response directions.
  • Publishable runs require an approved ARM64 host and a pre-run load average below 2.0.
  • Quick and partial matrices are rejected; failed scenarios remain visible with their captured errors.

Results compare only within one snapshot, host, workload, and benchmark layer. They are not universal hardware-independent rankings.