Profile
Back to NewsBack
Hacker News 13 min
Reader Mode
Compression and Concatenation on the Web

Compression and Concatenation on the Web

1 day ago
gzip

Compression and concatenation on the web

Everybody knows you can concatenate gzip files to form valid gzip files, but do the browsers know? We explore concatenation with gzip, brotli, and zstd.

You are getting early access to this article as a subscriber. Your support makes articles like this possible. Thank you.

The less data we send to the browser, the better. Compression is one of the ways we can send less data. But compression is not free. So we trade compression ratio (how much smaller can I get the input data) for speed of compression. Speed of decompression might also be a concern, but in browser-land it’s usually concluded that the quality of the compression (how many fewer bytes you can send over the wire) is more important than the speed of the decompression.

When data only comes fully together in response to a request (e.g. the response includes data for a specific user, pulled from a database), we might choose to compress the response dynamically in the server. In this scenario we want compression to be relatively cheap so that compression does not become a large part of response latency. Whereas when data is fully static (whether it’s images or CSS or JavaScript or even HTML), we can afford to do more expensive compression if it will produce smaller output. We do it once and benefit thereafter.

The IANA has a dozen registered content encoding schemes including gzip, brotli, and zstd. But HTTP Archive’s 2025 Web Almanac tells us only gzip, brotli, and zstd are in major use. And that brotli is now the plurality used among CDN traffic, while gzip is the majority among non-CDN traffic. gzip has been around forever so this makes sense. Brotli became widely available in 2017 while zstd has only just this year become available on (the latest version of) every major consumer platform. We’ll see how usage has changed in the upcoming 2026 Web Almanac, maybe brotli will have taken over non-CDN traffic.

Let’s take a look at the basics of how we might do compression, and some of the trouble we might run into when joining compressed files together in various formats.

Compressing at build time#

Say we have a web server.

import http.server


class Handler(http.server.SimpleHTTPRequestHandler):
  def do_GET(self):
    body = open(self.translate_path(self.path), "rb").read()
    self.send_response(200)
    self.end_headers()
    self.wfile.write(body)


http.server.test(HandlerClass=Handler)

server_nocompress.py

And some HTML.

<h1>Hello world!</h1>

index.html

We can run the server.

$ python3 server_nocompress.py
Serving HTTP on 0.0.0.0 port 8000 (http://0.0.0.0:8000/) ...

And in another window curl it.

$ { curl -sv http://localhost:8000/index.html; } 2>&1 | grep -v '^[*{}]'
> GET /index.html HTTP/1.1
> Host: localhost:8000
> User-Agent: curl/8.5.0
> Accept: */*
>
< HTTP/1.0 200 OK
< Server: SimpleHTTP/0.6 Python/3.12.3
< Date: Sun, 04 Oct 2026 21:15:40 GMT
<
<h1>Hello world!</h1>

Great. Let’s talk about compression for the web.

Compression for the web#

The goal of compression in the web environment is to send fewer bytes from server to client. But clients have to opt into this by passing an Accept-Encoding header with a list of compression formats they support (e.g. gzip, brotli, zstd).

We can compress static files dynamically when we see an Accept-Encoding header.

import gzip
import http.server


class Handler(http.server.SimpleHTTPRequestHandler):
  def do_GET(self):
    body = open(self.translate_path(self.path), "rb").read()
    self.send_response(200)
    if "gzip" in self.headers.get("Accept-Encoding", ""):
      body = gzip.compress(body)
      self.send_header("Content-Encoding", "gzip")
    self.end_headers()
    self.wfile.write(body)


http.server.test(HandlerClass=Handler)

server_dynamic_compress.py

Run the server.

$ python3 server_dynamic_compress.py
Serving HTTP on 0.0.0.0 port 8000 (http://0.0.0.0:8000/) ...

We can pass the Accept-Encoding header ourselves manually to get back the gzipped content, but we’d have to pass the body through gunzip to get back the uncompressed data.

$ { curl -sv -H "Accept-Encoding: gzip" http://localhost:8000/index.html | gunzip; } 2>&1 | grep -v '^[*{}]'
> GET /index.html HTTP/1.1
> Host: localhost:8000
> User-Agent: curl/8.5.0
> Accept: */*
> Accept-Encoding: gzip
>
< HTTP/1.0 200 OK
< Server: SimpleHTTP/0.6 Python/3.12.3
< Date: Sun, 04 Oct 2026 21:16:27 GMT
< Content-Encoding: gzip
<
<h1>Hello world!</h1>

curl has a shortcut for this intrinsic part of the web, --compressed.

$ { curl -sv --compressed http://localhost:8000/index.html; } 2>&1 | grep -v '^[*{}]'
> GET /index.html HTTP/1.1
> Host: localhost:8000
> User-Agent: curl/8.5.0
> Accept: */*
> Accept-Encoding: deflate, gzip, br, zstd
>
< HTTP/1.0 200 OK
< Server: SimpleHTTP/0.6 Python/3.12.3
< Date: Sun, 04 Oct 2026 21:17:26 GMT
< Content-Encoding: gzip
<
<h1>Hello world!</h1>

But back to our compression. We supported dynamic compression at runtime, but that’s inefficient since this file was static already. We could instead compress index.html into index.html.gz out of band, and instead of compressing index.html on the fly during a request, we instead send back the contents of index.html.gz.

import http.server


class Handler(http.server.SimpleHTTPRequestHandler):
  def do_GET(self):
    path = self.translate_path(self.path)
    self.send_response(200)
    if "gzip" in self.headers.get("Accept-Encoding", ""):
      path += ".gz"
      self.send_header("Content-Encoding", "gzip")
    self.end_headers()
    self.wfile.write(open(path, "rb").read())


http.server.test(HandlerClass=Handler)

server_static_compress.py

Run gzip and start the server.

$ cat index.html | gzip > index.html.gz
$ python3 server_static_compress.py
Serving HTTP on 0.0.0.0 port 8000 (http://0.0.0.0:8000/) ...

Then fetch the file.

$ { curl -sv --compressed http://localhost:8000/index.html; } 2>&1 | grep -v '^[*{}]'
> GET /index.html HTTP/1.1
> Host: localhost:8000
> User-Agent: curl/8.5.0
> Accept: */*
> Accept-Encoding: deflate, gzip, br, zstd
>
< HTTP/1.0 200 OK
< Server: SimpleHTTP/0.6 Python/3.12.3
< Date: Sun, 04 Oct 2026 21:23:10 GMT
< Content-Encoding: gzip
<
<h1>Hello world!</h1>

And now we’re able to serve both compressed and uncompressed data without any runtime penalty.

Concatenating compressed data#

Compression is expensive. You’d like to be able to do it once and not again. And sometimes you have components of a file that are built at different times, or built at different frequencies. You’d like to be able to join them together again to form a valid whole compressed stream without having to decompress the parts, concatenate them, and compress the result, where the compressed concatenated result can be read by any reasonable client.

gzip is famous for being concat-able.

echo '<title>My website</title>' | gzip > header.html.gz
echo '<h1>Hello world!</h1>' | gzip > body.html.gz
cat header.html.gz body.html.gz > index.html.gz

And you get the two inputs back.

$ gunzip -c index.html.gz
<title>My website</title>
<h1>Hello world!</h1>

But downstream applications (HTTP clients, web browsers, etc.) don’t necessarily handle this well or at all.

$ curl --compressed http://localhost:8000/index.html
<title>My website</title>
curl: (23) Failed writing received data to disk/application

Not lovely! But at least curl told you something went wrong. (And this message has gotten clearer this year too.)

Grab Chrome.

curl -O https://dl.google.com/linux/direct/google-chrome-stable_current_amd64.deb
sudo apt install -y ./google-chrome-stable_current_amd64.deb

And try it out.

$ google-chrome --headless --no-sandbox --dump-dom http://localhost:8000/index.html
[138026:138026:1004/214112.064341:ERROR:dbus/object_proxy.cc:572] Failed to call method: org.freedesktop.DBus.Properties.GetAll: object_path= /org/freedesktop/UPower/devices/DisplayDevice: org.freedesktop.DBus.Error.ServiceUnknown: The name org.freedesktop.UPower was not provided by any .service files
<html><head><title>My website</title>
</head><body></body></html>

Chrome won’t even tell you the second part was dropped.

So while gzip, the CLI, “supports” concatenation, other users might not. The zlib repository (which is another implementation of gzip, which most languages use when they support gzip if they do not implement the gzip format themselves, whereas gzip is purely an application) has an example program, gzjoin.c, for joining multiple gzip files. It does require decompressing every file, but the decompression is only for searching the file. It does not require recompression to complete the join.

Grab zlib itself and a C compiler and build gzjoin.c, whose source comes in the zlib package.

sudo apt update -y
sudo apt-get install -y build-essential zlib1g-dev
cc -o gzjoin /usr/share/doc/zlib1g-dev/examples/gzjoin.c -lz

Safely join our two gzipped files and curl them against the still-running Python server.

$ ./gzjoin header.html.gz body.html.gz > safeindex.html.gz
$ curl --compressed http://localhost:8000/safeindex.html
<title>My website</title>
<h1>Hello world!</h1>

And with Chrome.

$ google-chrome --headless --no-sandbox --dump-dom http://localhost:8000/safeindex.html
[138218:138218:1004/214243.315401:ERROR:dbus/object_proxy.cc:572] Failed to call method: org.freedesktop.DBus.Properties.GetAll: object_path= /org/freedesktop/UPower/devices/DisplayDevice: org.freedesktop.DBus.Error.ServiceUnknown: The name org.freedesktop.UPower was not provided by any .service files
<html><head><title>My website</title>
</head><body><h1>Hello world!</h1>
</body></html>

So gzjoin.c is a great example, but I do not think any major programming language’s implementation of gzip has an API for doing this. Of course any developer can reimplement gzjoin.c for their environment, but it would be nice to be built in.

Brotli#

The reference implementation of brotli (by Google) does not support gzip-style concatenation.

Grab Google brotli.

sudo apt install -y brotli

Try joining two brotli compressed files.

echo '<title>My website</title>' | brotli > header.html.br
echo '<h1>Hello world!</h1>' | brotli > body.html.br
cat header.html.br body.html.br > index.html.br

And the second part gets corrupted.

$ brotli --decompress --stdout header.html.br
<title>My website</title>
$ brotli --decompress --stdout body.html.br
<h1>Hello world!</h1>
$ brotli --decompress --stdout index.html.br
<title>My website</title>
corrupt input [index.html.br]

Dropbox developers tweaked compression slightly in 2018 (and wrote about it in 2020) in their rust-brotli implementation to support concatenation with a tool similar to gzjoin. And then in 2025 the rust-brotli folks also added flags for building brotli files that you can actually cat together without an external tool (you just have to remember to append a specific byte when the stream is done).

Grab their CLI.

sudo apt-get install -y cargo
cargo install brotli

And build the two compressed parts.

echo '<title>My website</title>' | ~/.cargo/bin/brotli -c -appendable -bytealign -bare > header.html.br
echo '<h1>Hello world!</h1>' | ~/.cargo/bin/brotli -c -catable -appendable -bytealign -bare > body.html.br
{ cat header.html.br body.html.br; printf '\x03'; } > index.html.br

And now even Google brotli can decode this.

$ ~/.cargo/bin/brotli index.html.br
<title>My website</title>
<h1>Hello world!</h1>
$ brotli --decompress --stdout index.html.br
<title>My website</title>
<h1>Hello world!</h1>

This offers superior flexibility to being able to join brotli streams by specifying stream offsets for each part you’ll later join, which is the closest Google brotli gets. And support for stream offsets in client libraries of Google brotli, even when kept in the same repo as Google brotli, is patchy. Stream offsets are not even accessible from the brotli CLI.

Still, both are a tad better supported than gzjoin.c which only exists as an example file and not even an API. And both options for brotli will produce compressed files that curl and browsers can read.

We can see this with another Python server that only handles brotli.

import http.server


class Handler(http.server.SimpleHTTPRequestHandler):
  def do_GET(self):
    path = self.translate_path(self.path)
    self.send_response(200)
    if "br" in self.headers.get("Accept-Encoding", ""):
      path += ".br"
      self.send_header("Content-Encoding", "br")
    self.end_headers()
    self.wfile.write(open(path, "rb").read())


http.server.test(HandlerClass=Handler)

server_static_compress_br.py

Run it.

$ python3 server_static_compress_br.py
Serving HTTP on 0.0.0.0 port 8000 (http://0.0.0.0:8000/) ...

And it works with curl and Chrome.

$ curl --compressed http://localhost:8000/index.html
<title>My website</title>
<h1>Hello world!</h1>
$ google-chrome --headless --no-sandbox --dump-dom http://localhost:8000/index.html
[139661:139661:1004/231920.832523:ERROR:dbus/object_proxy.cc:572] Failed to call method: org.freedesktop.DBus.Properties.GetAll: object_path= /org/freedesktop/UPower/devices/DisplayDevice: org.freedesktop.DBus.Error.ServiceUnknown: The name org.freedesktop.UPower was not provided by any .service files
<html><head><title>My website</title>
</head><body><h1>Hello world!</h1>
</body></html>

Looks great.

zstd#

zstd was the only format built from the beginning to be concatenable on the web. Grab it.

sudo apt install -y zstd

Build our multi-part file.

echo '<title>My website</title>' | zstd > header.html.zst
echo '<h1>Hello world!</h1>' | zstd > body.html.zst
cat header.html.zst body.html.zst > index.html.zst

And it decompresses.

$ zstd --decompress -c index.html.zst
<title>My website</title>
<h1>Hello world!</h1>

Let’s try it out from a server.

import http.server


class Handler(http.server.SimpleHTTPRequestHandler):
  def do_GET(self):
    path = self.translate_path(self.path)
    self.send_response(200)
    if "zstd" in self.headers.get("Accept-Encoding", ""):
      path += ".zst"
      self.send_header("Content-Encoding", "zstd")
    self.end_headers()
    self.wfile.write(open(path, "rb").read())


http.server.test(HandlerClass=Handler)

server_static_compress_zstd.py

Run it.

$ python3 server_static_compress_zstd.py
Serving HTTP on 0.0.0.0 port 8000 (http://0.0.0.0:8000/) ...

And curl and the browser support it.

$ curl --compressed http://localhost:8000/index.html
<title>My website</title>
<h1>Hello world!</h1>
$ google-chrome --headless --no-sandbox --dump-dom http://localhost:8000/index.html
[139953:139953:1004/232250.964510:ERROR:dbus/object_proxy.cc:572] Failed to call method: org.freedesktop.DBus.Properties.GetAll: object_path= /org/freedesktop/UPower/devices/DisplayDevice: org.freedesktop.DBus.Error.ServiceUnknown: The name org.freedesktop.UPower was not provided by any .service files
<html><head><title>My website</title>
</head><body><h1>Hello world!</h1>
</body></html>

Kudos to zstd making this so easy!

The evolving specification#

The interesting thing about this is not that there are shortcomings in a specification but that implementations like rust-brotli were able to work around them without building files that other implementations could not decode. Tweaking a standard is common elsewhere, such as implementations of consensus protocols.

Noticed a mistake? Have a question or comment? Write to the editor.

Comments

Chat with me