BlogEngineering
Autoscaling workers from one Postgres queue on Railway
Notes from moving our transcriptions off the web server onto workers that share a PostgreSQL queue and size themselves on Railway, including the API call that failed first.
The short answer
On October 7, 2026 we split ScoreStarling’s one Railway service in two: the site only serves requests, and one to four worker replicas run the transcriptions. The queue stayed a table in our PostgreSQL database: a worker claims the oldest job under a transaction-level advisory lock and reports every 15 seconds, and another worker takes back any job that goes silent for two minutes. Railway doesn’t add replicas by itself, so the leading worker sets the count from the queue through Railway’s API, after a first version that asked for a field the API doesn’t have.
Why move transcription off the web server?
Because the web server was also the only transcription worker, and it could grow only by getting a bigger container. Until October 7, one Railway replica with 2 vCPUs and 8 GB served the site, the API and the MCP endpoint, and ran every transcription, one at a time. Only the holder of a session-level advisory lock worked the queue, so a second replica would have added no transcription capacity.
Now the same image runs as two Railway services with a role setting: web never takes a job, and worker has no public address. A job is one transcription (ScoreStarling turns recordings into sheet music): 17 to 266 seconds of a worker’s time in the production runs we sampled from October 2 to 7.
Why keep the queue in Postgres instead of a broker?
Because the job row already holds what a broker would only point to: owner, status, credit reservation, price and files. A broker would be a second store to keep in step, and most brokers deliver at least once, while a band transcription is a paid request to an outside service that must never be sent twice. We would need our own ownership rules anyway.
The load is small: a few dozen waiting jobs at most, each taking seconds to minutes. Railway’s queue guide says a Postgres queue has fewer moving parts when you already run Postgres, at lower throughput than Redis. Queue libraries such as pg-boss would add a second job table. We’ll reconsider at tens of jobs a second.
How do several workers share one Postgres queue without taking the same job?
Each claim is one short transaction under one lock for the whole queue. An idle worker checks without the lock whether anything is queued, and if not, looks again two seconds later:
-- One claim, slightly simplified
SELECT 1 FROM jobs WHERE status = 'queued' LIMIT 1; -- nothing queued: stop here
SELECT pg_advisory_xact_lock(<queue key>); -- held until COMMIT
SELECT * FROM jobs WHERE status = 'queued' ORDER BY created_at LIMIT 1;
UPDATE jobs SET status = 'running', worker_id = $1, heartbeat_at = $2
WHERE id = $3;
COMMIT;
A second worker waits at the lock, and its next query starts after the first has committed. At Read Committed, PostgreSQL’s default, a query sees what was committed before it began, so the job already shows as running. In our checks, six workers on separate connections drained 40 queued jobs, each taken exactly once.
With many workers, SELECT … FOR UPDATE SKIP LOCKED is the usual choice; PostgreSQL’s documentation suggests it for queue-like tables. We kept one lock because uploads, credit changes and recoveries already take it (in the local SQLite version it is BEGIN IMMEDIATE), so claims take turns with them. We haven’t measured that wait. The queue also needs no database session of its own: no LISTEN, which PgBouncer’s transaction pooling doesn’t support, and no session-level lock like the one our old single leader held.
What happens to a job when its worker stops?
- QueuedThe site saves the upload and adds a job row.
- ClaimedA free worker takes the oldest job under the lock.
- RunningThe worker reports every 15 s, with the stage reached.
- FinishedThe score is saved, and the worker looks for more.
It is recovered, and paid work is never sent twice. A worker stopped on purpose, by a deploy for example, puts its free job back in the queue itself. A worker that crashes just goes silent. Every 15 seconds a worker writes the time and its stage to its job’s row, and every 30 seconds each worker looks, under the same lock, for running jobs whose 120-second lease has run out without a report. A free transcription goes back to the queue once; a second interruption fails it. A band transcription resumes from its saved request, or waits for review without one. A worker that was only slow, and finds its job given away, kills its process and writes nothing.
Once-only work, such as the controller, runs on the leader: the lowest ID among workers that reported to a workers table in the last 45 seconds, which each does every 15. Nothing is held, so a dead leader is replaced within 45 seconds.
Why scale workers on queue depth instead of CPU?
Because a worker runs one job at a time, the workers needed equal the jobs queued or running. A busy worker’s CPU says it is busy, not whether one job or twenty are waiting, and Railway’s autoscaling guide names queue depth as the signal for workers too. Railway grows a container up to its CPU and memory limits by itself, but a replica count stays where you set it; the guide’s controller scales up at once and down a step at a time.
Ours cuts back differently. The API takes a count and Railway picks which replica to stop, so we cut back only when nothing is queued or running and no worker is busy, then straight to one.
| Rule | Value | Why |
|---|---|---|
| Time between looks at the queue | 30 s | Only the leader looks, so each decision is made once. |
| Workers kept, and the most allowed | 1 to 4 | One is ready for the next job; the cap bounds the bill. |
| Wait before a job asks for a worker while one looks free | 30 s | A free worker normally takes it within 2 s; with none free, the next look asks. |
| Time between two increases | 90 s | New workers get time to start before the next ask. |
| Time a new worker is given to appear | 300 s | Until then the controller doesn’t ask again. |
| Quiet time before cutting back to one | 600 s | With no job and no busy worker, the replica Railway stops is idle. |
When jobs wait, the controller asks for one worker per queued or running job, at least one more than it has and at most four.
What broke in production?
The controller’s first call to Railway’s real API; our tests had used a stand-in that answered any query. Five minutes after the change that switched the controller on was merged, every look at the queue failed. The workers kept taking jobs, but the log and Sentry said only RailwayError: our logs leave out exception messages.
The same query sent without a token reproduced it. Railway validated it against its schema first and answered HTTP 400, GRAPHQL_VALIDATION_FAILED. Our service is placed by region, so the write sets multiRegionConfig, which the update input accepts, but our read asked a service instance for that field, and it has none. The fix reads the count from the environment’s config, which our Railway config file and railway scale also change.
# Before: rejected with GRAPHQL_VALIDATION_FAILED
query ($service: String!, $environment: String!) {
serviceInstance(serviceId: $service, environmentId: $environment) {
numReplicas multiRegionConfig
}
}
# After: the count is at services.<service id>.deploy.multiRegionConfig
query ($environment: String!) {
environment(id: $environment) { config(decryptVariables: false) }
}
# The write, unchanged; input: {"multiRegionConfig": {"<region>": {"numReplicas": 2}}}
mutation ($service: String!, $environment: String!, $input: ServiceInstanceUpdateInput!) {
serviceInstanceUpdate(serviceId: $service, environmentId: $environment, input: $input)
}
The log now carries Railway’s own message, which never contains the token, since a refused token and a refused query were both HTTP 400. The stand-in refuses the old query. And a new check sends both operations to the live schema with placeholder IDs and no token: GraphQL servers validate before executing, so a wrong field fails validation, while a valid query gets as far as Railway’s “Not Authorized”.
The queue may hold several times more waiting jobs, still behind one worker.
Merged: workers share the queue, and a worker service starts beside the site. In a test, two uploads run at once.
Merged: the site stops taking jobs, and the controller is switched on.
Every look at the queue fails with
RailwayError. Jobs still run.Merged: the count is read from the environment’s config.
The first count read from Railway: one worker, steady.
Two jobs at once. The controller asks for a second worker, which is up 17 seconds later.
After 602 seconds with nothing to do, back to one. No errors through 06:28.
The same day, the leader began saving each look’s outcome for the site’s health page, which warns when Railway refuses the token or no look arrives for two minutes. The site’s own health check no longer covers the workers, so a worker outage now shows in its five-minute service check, and Sentry emails us when a service stays unwell twice in a row.
What we haven’t verified, and what it costs
- Only free transcriptions have run on the workers; a band transcription there, and its recovery after a lost worker, are untested.
- Our test account may have two scores in progress, and we skipped the planned days of dry-run logging, so the controller has met one small burst and an hour of quiet. Three or four workers, the 90-second step and stopping a busy replica haven’t happened in production.
- Production doesn’t record when a job is taken, so we can’t say yet how much shorter the wait in line is. In the scale-up test the new worker found nothing to do: the first had already taken the waiting job.
- After ten quiet minutes the pool is back to one, so a burst waits up to 30 seconds for the next look, then for a new worker to start: 17 seconds in our one test.
- At the maximum, four jobs run and the rest wait in line.
- A deploy restarts every worker, returning a running free transcription to the queue once, and each CI run that applies our Railway config resets the count to one until the next look.
Railway charges $20 per vCPU and $10 per GB of memory a month, by the minute (checked October 7, 2026). We estimate a worker at about $0.08 an hour while it transcribes at 2 vCPUs and 1.5 GB, and much less while it waits. A hard usage limit on Railway takes every workload offline when reached, so raise such a limit together with the maximum.
If you build this yourself
- Keep the queue with the job’s state unless you need a broker’s routing or throughput.
- Claim in one short transaction; with many workers, use
SKIP LOCKED. - Keep session state out of the queue: no
LISTEN, no session-level locks. - Give each running job an owner and a heartbeat, and make a worker that lost its job stop without writing. Never resend a paid outside request; resume it from its saved ID.
- Scale on queued plus running jobs, and cut back only when no worker is busy if the platform picks which replica stops.
- Test the platform’s API against its live schema, and log its error messages once you know they hold no secrets.
- Give the controller somewhere to report.
Sources
- 13.3. Explicit Locking: 13.3.5 Advisory Locks, PostgreSQL 18 documentation — PostgreSQL Global Development Group
- 9.28. System Administration Functions: 9.28.10 Advisory Lock Functions, PostgreSQL 18 documentation — PostgreSQL Global Development Group
- SELECT: The Locking Clause, PostgreSQL 18 documentation — PostgreSQL Global Development Group
- 13.2. Transaction Isolation: 13.2.1 Read Committed Isolation Level, PostgreSQL 18 documentation — PostgreSQL Global Development Group
- PgBouncer features — PgBouncer
- Autoscale a Service Horizontally Based on Load — Railway
- Choose Between Cron Jobs, Background Workers, and Queues — Railway
- Pricing Plans — Railway
- Cost Control — Railway
- GraphQL Specification, October 2021 Edition: 6.1.1 Validating Requests — GraphQL Foundation
Questions and answers
Does Railway autoscale replicas horizontally?
Not as of October 7, 2026. Railway gives a container more CPU and memory up to its limits, but a service’s replica count stays where you set it. Its autoscaling guide has you run a controller that reads a load signal and calls the API’s serviceInstanceUpdate.
Should a Postgres job queue use SKIP LOCKED or an advisory lock?
FOR UPDATE SKIP LOCKED lets many workers claim different rows at once, and PostgreSQL’s documentation suggests it for queue-like tables. A single transaction-level advisory lock makes claims take turns, which is simpler for a few workers, especially when the same lock already guards your other writes.
Why does Railway’s API answer GRAPHQL_VALIDATION_FAILED for multiRegionConfig?
In our case, because we asked a service instance to return multiRegionConfig, which only the update input has. Read the count from environment(id) { config } at services.<service id>.deploy.multiRegionConfig, and write it with serviceInstanceUpdate.
What happens to a running job when a worker replica is removed?
Our controller removes replicas only while no worker is busy. A busy worker stopped anyway, by a deploy for example, returns a free transcription to the queue once; if it vanishes, another worker takes the job back after 120 seconds without a report. Paid requests are never sent again.
How much does an extra worker cost on Railway?
Railway bills $20 per vCPU and $10 per GB of memory a month, by the minute (checked October 7, 2026). We estimate about $0.08 an hour for a worker transcribing at 2 vCPUs and 1.5 GB, and much less while it waits. Our cap is four workers.