Shahrah is a PostgreSQL proxy written in Rust for geo-distributed PostgreSQL deployments.

A PostgreSQL proxy that knows which region a user's data is in.

It shards by a key you name, splits reads from writes, pools connections, and — the part nothing else does — keeps a directory of which region each key lives in, routes to that region, and moves a key between regions while the application keeps reading and writing.

Applications talk to it in the PostgreSQL wire protocol. There is no driver to install and no shahrah-specific code in the request path.

What it is for

Hundreds of millions of users, spread over regions, where a user's data should be in the region they are in. Sharding alone puts a Tehran user's rows on whichever shard the hash picks, and if that shard is in Virginia, every read crosses an ocean. shahrah puts the region in front of the shard map rather than inside it, so the shard a key lives on is chosen within its home region.

Measured on a three-region bed with the real latencies shaped in, a user's read:

data in the user's region one region for everyone
a user in na-east 0.15 ms 0.17 ms
a user in eu-central 0.20 ms 43.22 ms
a user in asia-west 0.16 ms 76.16 ms

Five rounds, 100 reads per region per arm. The na-east row is the honest cost: a directory lookup that a single-region deployment does not have to do. See docs/NUMBERS.md for the spreads, and for where shahrah is slower than the alternatives.

Running it against three shards

Three PostgreSQL servers, one topology file, one process.

# shards.toml
region = "na-east"

[policy.keys.users]
column = "id"
type = "int"

[[shards]]
number = 1
region = "na-east"
primary  = { address = "10.0.0.11:5432", region = "na-east" }
replicas = [{ address = "10.0.0.12:5432", region = "na-east" }]
logical  = [{ start = 0, end = 21844 }]

[[shards]]
number = 2
region = "na-east"
primary = { address = "10.0.0.21:5432", region = "na-east" }
logical = [{ start = 21845, end = 43689 }]

[[shards]]
number = 3
region = "na-east"
primary = { address = "10.0.0.31:5432", region = "na-east" }
logical = [{ start = 43690, end = 65535 }]

Every shard needs a role shahrah can log in as, and a function it can ask for a user's stored verifier — shahrah never sees a password in the clear, and neither do you:

create role shahrah login superuser password 'change-me';
create or replace function shahrah_get_auth(in wanted text,
                                            out username text, out verifier text)
returns record as $$
  select rolname::text, rolpassword::text from pg_authid where rolname = $1
$$ language sql security definer;
revoke all on function shahrah_get_auth(text) from public;

Then:

SHAHRAH_LISTEN=0.0.0.0:6432 \
SHAHRAH_CONFIG=/etc/shahrah/shards.toml \
SHAHRAH_BACKEND=10.0.0.11:5432 \
SHAHRAH_BACKEND_USER=shahrah \
SHAHRAH_BACKEND_PASSWORD=change-me \
SHAHRAH_METRICS_LISTEN=127.0.0.1:9187 \
  shahrah-proxy

Point the application at postgresql://youruser:yourpassword@proxy:6432/yourdb. Nothing else changes.

The settings that matter

variable what it does
SHAHRAH_LISTEN where clients connect.
SHAHRAH_CONFIG the topology file. Without one, every statement goes to SHAHRAH_BACKEND unchanged.
SHAHRAH_BACKEND where shahrah reads the auth verifier, and the fallback backend.
SHAHRAH_BACKEND_USER, SHAHRAH_BACKEND_PASSWORD how shahrah logs in to the shards.
SHAHRAH_MAX_POOL connections per endpoint, per database, per role.
SHAHRAH_METRICS_LISTEN serves the live dashboard at / and /metrics, /snapshot, /key/<table>/<key>. Unset by default.
SHAHRAH_METRICS_PASSWORD the operator password. Without it shahrah refuses to bind anything but loopback, because the dashboard publishes the topology and where every user's data lives.
SHAHRAH_METRICS_OPEN yes to serve a reachable address with no password anyway, for a deployment that has put its own authentication in front.
SHAHRAH_METRICS_TLS_CERT, SHAHRAH_METRICS_TLS_KEY serve the dashboard over TLS. Without both it is plain HTTP, and shahrah says so when a password would cross a network in the clear.
SHAHRAH_WARM_PER_SHARD connections opened at start so no session pays the first crossing.
SHAHRAH_TRACE_SAMPLE trace one statement in N.
SHAHRAH_AUTH_DATABASE which database the verifier is read from. postgres unless you say otherwise.
SHAHRAH_AUTH_QUERY the statement that reads it, if shahrah_get_auth is not what you want to call it.

TLS

variable what it does
SHAHRAH_TLS_CERT, SHAHRAH_TLS_KEY serve TLS to clients. Without both, shahrah answers SSLRequest with N and the connection continues in the clear.
SHAHRAH_BACKEND_TLS require to insist on TLS to the shards. Anything else, including unset, means disable -- so this is opt-in, and a typo is silent.
SHAHRAH_BACKEND_TLS_INSECURE accept a shard certificate shahrah cannot verify. For a test bed; not for anything else.

The directory

variable what it does
SHAHRAH_DIRECTORY_TABLE the table each shard keeps the home regions in. shahrah_directory.
SHAHRAH_DIRECTORY_CACHE how many home regions a proxy holds in memory. A million.
SHAHRAH_DIRECTORY_TTL how long it trusts one before reading it again. 300 seconds. A move announced through the group invalidates it immediately, so this is the bound for a move shahrah was not told about.

More than one proxy

A single proxy needs none of these. They are how a group of them agrees on the topology and hears about a key that moved.

variable what it does
SHAHRAH_RAFT_ID this proxy's number in the group.
SHAHRAH_RAFT_LISTEN where its peers reach it.
SHAHRAH_RAFT_PEERS 1=host:port,2=host:port,..., every member including this one.
SHAHRAH_RAFT_STATE the file it keeps its log and state machine in.
SHAHRAH_RAFT_BOOTSTRAP set on exactly one proxy, exactly once, to create the group.
SHAHRAH_RAFT_SPREAD how the group carries the topology between members.

Connections

variable what it does
SHAHRAH_RESET_QUERY run on a backend connection before it goes back in the pool.
SHAHRAH_BACKEND_TIMEZONE the timezone shahrah sets on a backend connection.

Placing a table

A table is one of two things, and shahrah refuses to guess:

[policy.keys.users]      # geo-partitioned: rows live in one region, by this key
column = "id"
type = "int"             # int | text | uuid

[policy.replicated.plans]  # globally replicated: the same rows everywhere
writer_region = "na-east"  # written in one region, read from any

A statement against a geo-partitioned table that does not carry its key is refused rather than sent somewhere hopeful. If the key is there but shahrah cannot see it — a view, a join, an ORM you do not control — say so in a comment:

/* shahrah: key=7 */ select * from orders_view limit 50

Asking where a user's data is

The console is a database called shahrah, spoken over the same wire protocol, so psql reaches it:

psql -h proxy -p 6432 -U youruser shahrah -c 'WHERE IS users 100005'
 answer     | learned_from                 | logical | shard | endpoint       | local
 eu-central | the directory, read just now |    6655 |     3 | 10.1.0.11:5432 | no

SHOW HELP lists the rest: SHOW POOLS, SHOW HEALTH, SHOW TRAFFIC, SHOW ALERTS, SHOW TOPOLOGY, SHOW PLACEMENT, SHOW DIRECTORY, SHOW MOVERS, SHOW FLEET <subject> for every proxy in the group at once, and the verbs DRAIN, UNDRAIN, RELOCATE, REBALANCE, REPAIR, TRACE.

The console answers the simple query protocol. psql is fine; a driver that prepares every statement gets a clear error rather than rows, and should read /metrics instead.

Moving a user to another region

psql -h proxy -p 6432 -U youruser shahrah -c 'RELOCATE users 100005 TO eu-central'

shahrah takes an intent marker on the key's directory row by compare-and-set — so two operators racing produce one move and one refusal — copies the rows, switches every region's copy of the directory, waits for the other proxies to hear of it, and only then removes the rows it left behind. Statements for that key are held for the moment of the cutover rather than answered from the wrong place.

If a move is interrupted, REPAIR finishes it. If rows sit on the wrong shard of the right region, REBALANCE <region> moves them.

Watching it

/metrics is Prometheus text format: statements, refusals, per-endpoint traffic and errors, pool occupancy, replica lag measured against the primary, and one shahrah_alert series per condition worth waking someone for.

/ is the dashboard, and it is live: it opens a WebSocket to /live and the proxy pushes a snapshot every second, so the numbers move without the page reloading under whoever is reading it. The first frame carries the last three minutes of history so the charts are drawn before the first tick; every frame after it carries one new sample, which is about 1.6 KB a second per viewer rather than the 44 KB it would take to resend the history each time.

It lays out every database in every region with its state, role, replication lag and traffic; rates over time for statements routed, statements reaching a shard, statements crossing a region, refusals, backend errors, cache and directory hits, and pool occupancy; a diagram of every region with its shards and their endpoints, coloured by what the prober last saw; the conditions currently worth waking someone for; what every other proxy in the group sees, named individually, with an unreachable one named too; and a box to look up one user. It is built from the fleet view rather than from this proxy's own, so any proxy shows the whole world.

One page, no framework, no fonts or scripts fetched from anywhere: the charts are SVG the page draws itself. If a proxy between the operator and shahrah refuses to carry a WebSocket, the page notices after three attempts and falls back to polling /snapshot, saying so in the corner rather than sitting there looking current while it is not.

/snapshot is that same snapshot as JSON, for anything that would rather poll. /plain is the dashboard without JavaScript — server-rendered, no live updates — for a browser that will not run any.

/key/<table>/<key> is the page behind that box, and the one question no generic dashboard can answer: where this user's data is, which shard holds it in each region, and whether this proxy is the one nearest to it.

Who may look

The metrics port shows your topology, your replica lag, and which region any named user's rows are in. It asks for a password before it shows any of it.

Set SHAHRAH_METRICS_PASSWORD to something generated rather than something chosen:

openssl rand -base64 32

It is compared as a constant-time equality of its SHA-256 — enough to leak nothing by timing, but it is not a slow password hash, so a memorable password would not survive its digest getting out. Generate it, keep it wherever you keep secrets, and do not put it in a file you commit.

Without one, shahrah serves loopback with a warning and refuses to bind an address other machines can reach at all — there is no configuration in which it quietly listens to the network with nothing in front of it. SHAHRAH_METRICS_OPEN=yes overrides that for a deployment where something else already authenticates.

A browser signs in at /login and gets a session cookie: 32 random bytes, HttpOnly so no script can read it, SameSite=Strict so no other site can spend it, and twelve hours long. /logout ends it. Anything that is not a browser — Prometheus, curl — sends the same password as Authorization: Bearer <password> or HTTP Basic.

Guessing is bounded. Five wrong answers from one address and that address waits 30 seconds, then a minute, then two, up to fifteen; the right password does not open the lock early. Every wrong answer costs the guesser a quarter of a second whichever way it arrived, so a header is not a faster door than the form. Any one address is held to 600 requests a minute across every path. The table of addresses is itself bounded, because a table that grows with the attacker is the denial of service it was meant to prevent.

The password is compared by constant-time equality of its SHA-256, so a wrong guess takes the same time whatever it got right.

Set SHAHRAH_METRICS_TLS_CERT and SHAHRAH_METRICS_TLS_KEY and the port speaks TLS: the dashboard over https, the live feed over wss, and the session cookie gains Secure so a browser will not send it back over anything else. A cleartext request to that port is not answered at all rather than downgraded. Without both variables it is plain HTTP, and if a password would then be crossing an address other machines can reach, shahrah says so at startup rather than leaving you to assume otherwise.

The certificate is yours to provide and yours to rotate — shahrah reads it at startup and does not watch the file.

What this still does not do: there is one password, not an account per operator, so it says someone is an operator and not which one, and there is no audit of who looked at what. RELOCATE and DRAIN remain on the console rather than the dashboard, so the page reads and does not act.

Building

cargo build --release
cargo test --workspace

The workspace denies unsafe, unwrap, expect, slicing that can panic, and arithmetic that can overflow. There are no comments in the source: the names and the tests are meant to carry it.

Keywords:

  • PostgreSQL proxy
  • PostgreSQL sharding
  • PostgreSQL connection pooling
  • PostgreSQL read/write splitting
  • PostgreSQL multi-region
  • PostgreSQL geo-sharding
  • PostgreSQL routing
  • Rust PostgreSQL proxy
  • distributed PostgreSQL

Benchmarks

bench/ brings up shahrah, pgbouncer, pgcat and PgDog on one network, refuses to report a number until they are placed identically, and reports a distribution rather than a figure. bench/geo/ is the three-region bed the table above came from. See bench/README.md.