Profile
Back to NewsBack
Dev.to 6 min
Reader Mode
Running a BFT blockchain as one native binary on a 2 GB VPS — what broke and what held

Running a BFT blockchain as one native binary on a 2 GB VPS — what broke and what held

13 hours ago

I run a small Layer-1 called Janzeer: stake-free BFT, 15-second slots, two voting rounds, finality in one block. Mainnet went live on 2026-10-01. The whole network — four validator nodes, PostgreSQL, and a Telegram mini-game that pays players — runs on one 2 GB, single-core VPS. This post is about how that is possible and what went wrong on the way. No token talk; the coin is not on any exchange and that is not what this is about.

The constraint

The node is a Spring Boot application in Kotlin: Netty for P2P, JPA on PostgreSQL, a REST and JSON-RPC API, and it serves the explorer SPA. Four of those as JVMs would never fit in 2 GB with a database next to them. So every process is a GraalVM native image: a ~250 MB binary that boots in 3 seconds and is capped at -Xmx256m. The binary's pages are shared by the four processes, so the cgroup charge is about 200 MB per node. The game backend is a native image too. Peak memory of the whole box during a 6-hour rehearsal, with load: 933 MB.

Three rules keep the native image working, and each one was learned the hard way:

  1. No @Lazy injections. A lazy CGLIB proxy cannot be generated in the image. ObjectProvider does the same job.
  2. Every class that is instantiated reflectively must be registered. The wire codec builds messages by reflection. The tracing agent only sees the messages that happen to flow during the trace, so the hints are written by hand for every message type. The first image pushed no WebSocket notifications because the notification envelope was missing from the hints — Jackson serialized it fine on the JVM and silently failed in the image.
  3. Hibernate's bytecode provider off (hibernate.bytecode.provider=none) and a capped DB pool. An image otherwise grows towards a share of physical RAM per process.

The 7-day acceptance test

Before going live, the public test network ran for a week with a monitor that scanned the logs for consensus errors, killed and relaunched nodes, and pushed load. Thirty-seven findings, all fixed before launch. The ones worth retelling:

Four nodes booting together all hit pool.ntp.org in the same instant. Two got no answer and waited the full 60-second retry before asking again. Slot ownership is clock-based, so a node without a trusted clock does not produce: 75 seconds of empty slots at every joint start. The fix is boring and effective: a jittered exponential retry, 5 s → 10 s → 20 s, capped at 60.

A node relaunched on its own database started two sync sessions at once when two peers answered its availability request, and committed a duplicated epoch genesis. Sync is now single-flight on one thread: the first answer starts the session, later answers are ignored, and a uniqueness constraint on block hash, height and epoch makes the duplicate impossible at the database level.

The periodic tip check, when it ran on validators, kept a stall alive forever. The check flips the node to PROCESSING, and a PROCESSING node drops every vote. Four anchors doing that in lock-step during a stall could never reach quorum again. Validators now catch up only through the live block that proves they are behind; the periodic check runs on watch-only nodes.

PostgreSQL enforces read-only transactions; H2 does not. The JSON-RPC send* methods ran inside a read-only transaction template. Every test passed on H2. PostgreSQL refused the first real write. Anything that touches transaction boundaries is now tried on PostgreSQL before it ships, and a test pins the transaction attributes of those methods.

The first live upgrades

Shipping to the box is one script: it stages the native binaries, launchers, the game and the website into a folder, verifies that no secret ever enters it, and a second script on the VPS installs files, renders systemd units and nginx vhosts, and restarts what changed.

The first upgrades restarted all four validators together. That cost about a minute of no blocks at first, then four minutes last week. The log showed why: at boot the node re-validated the last two epochs on disk — 17,000 blocks per epoch, 2.5 minutes on one core — while reporting itself as PROCESSING and not dialing anyone. Every restart paid that.

Two changes fixed it:

  • A clean-shutdown marker. At @PreDestroy the node writes the tip hash next to its config file. At boot it reads and deletes the marker; if the stored tip still matches, the blocks on disk were validated when they were committed and nothing changed since, so the sweep is skipped. A crash, an older build or a moved tip leaves no matching marker and the sweep runs as before. Restart time went from 2.5 minutes to 11 seconds.
  • Rolling restarts. The quorum is 3 of 4, so one node may be down. Slot ownership is round-robin, so a node restarted right after its own block has 60 seconds before its next turn. The deploy script now restarts one validator at a time, each right after it produced, and moves on only when it is back, synchronized and at the tip. It falls back to a joint restart only when the protocol version changed, because a node refuses a peer that speaks another version. The last upgrade: four nodes rolled, the chain never stopped, largest gap between blocks 15 seconds — one slot.

A detail that bit us once: right after a restart, the other nodes can report one peer fewer for a few seconds, because a fresh connection is registered on one side before the other. The first production roll stopped before the fourth node on that. The pre-check now waits up to 90 seconds for the network to settle.

What I would tell someone doing the same

  • The native image is worth it for small boxes, but treat the tracing agent's output as a supplement. Write the reflection hints for everything your codec or your JSON envelope touches.
  • Test your database semantics on the database you deploy. H2 in PostgreSQL mode is not PostgreSQL.
  • Measure boot. A node that is "up" in 3 seconds but busy for 3 minutes is down for 3 minutes.
  • Never restart a quorum together if you can avoid it. With round-robin slots and a 2f+1 quorum, a rolling restart timed after each node's own block costs zero blocks.
  • An alert bot that reports the symptoms of your own planned restart trains you to ignore it. Mute the transient symptoms for the planned window, never the real dangers.

The explorer has an animated page of the whole cycle, driven by the live chain, if you want to see the slots and votes move: https://explorer.janzeer.org/workflow. The SDKs (TypeScript, Dart, Kotlin, Python, Go) and the wallet are open source at https://github.com/janzeerorg; the node is a public download with a closed source for now.

Chat with me