🔒 The Silent Freeze

How a pipe, a lock, and a slow consumer deadlocked a crawler — and how to fix it

The Mystery

A production crawler froze mid-run: the process was alive, using ~0% CPU, but producing no logs, no errors, and no exit.

❄️ Frozen at 0% CPU

Not a crash — the process was blocked, waiting for something that would never finish.

🤫 No Log Output

Progress logs stopped because the print came after a blocking write().

🔁 Concurrency Made It Worse

One stalled worker turned into a total freeze when every other thread waited behind it.

🎯 A Real Bug

This happened in the Indeed crawler: crawler | feeder, where the feeder did slow S3/SQS/MySQL work.

The Pipeline That Froze

👷 5 Crawler Workers
➔
🔒 Shared Lock
➔
🗃️ 64 KB Pipe
➔
📥 Feeder (S3/SQS/MySQL)

Workers write NDJSON to stdout (a pipe). The feeder reads it and does slow I/O in between.

The Computer Science Behind It

🔧 Producer-Consumer Problem

A classic concurrency pattern: producers generate data, consumers process it, and a bounded buffer connects them.

When the consumer is slower than the producers, the buffer fills — this is called backpressure.

🔒 Mutex / Lock

A lock guarantees mutual exclusion: only one thread enters a critical section at a time.

Rule of thumb: never hold a lock across blocking I/O, because the lock then blocks everyone else too.

⏳ Blocking I/O

write() to a full pipe does not return until the reader drains space. The OS gives you no timeout and no error — it just blocks.

🔄 Deadlock vs Starvation

This bug is a form of deadlock: one thread holds a lock while blocked on I/O, and the rest wait on the lock — nobody can proceed.

The Deadlock Chain

1 Feeder pauses (slow S3/SQS)
→
2 Pipe buffer fills (64 KB)
→
3 write() blocks while holding the lock
→
4 All other workers block on the lock

One blocking write() under a shared lock = a frozen thread pool.

✅ Safe Pattern

data = build_record()
queue.put(data)      # fast, non-blocking

Producers enqueue and return. One dedicated thread does the I/O.

❌ Unsafe Pattern

with lock:
    write(data)      # blocks if full!

Holding the lock across write() couples every worker to the slowest I/O.

The Bug

Every job record was written under one shared lock:

❌ Before (buggy)

record = self.extract_job(job, query, location, crawl_id)
line = json.dumps(record, ensure_ascii=False) + "\n"
with self._output_lock:
    for out_f in outputs:
        out_f.write(line)   # blocks if pipe is full
        out_f.flush()
total_written += 1

Watch It Freeze

🔒 LOCK
📥 Feeder
📢 Event Log

Reproduce It — Minimum POC

The bug is not specific to crawlers. Here is the smallest faithful reproduction.

Minimal Python POC

import threading
import time

lock = threading.Lock()
pipe = []                     # stand-in for the 64 KB OS pipe

def producer():
    for i in range(50):
        with lock:             # BUG: lock held across a blocking write
            while len(pipe) >= 20:
                time.sleep(0.01)   # pretend write() blocks on a full pipe
            pipe.append("job-" + str(i))
        time.sleep(0.05)

def consumer():
    while True:
        if pipe:
            pipe.pop(0)
        time.sleep(0.15)       # slow consumer -> backpressure

threading.Thread(target=consumer, daemon=True).start()
producer()

Run this and the producer eventually stalls inside the lock while the slow consumer drains — the classic freeze.

Interactive Simulator

🔒
📥 Consumer
📢 Simulator Log

The Fix

Move the blocking I/O off the worker threads onto one dedicated writer thread fed by a bounded queue.

✅ After (fixed)

class NdjsonWriter:
    def __init__(self, streams, maxsize=20000):
        self.queue = queue.Queue(maxsize=maxsize)
        self._thread = threading.Thread(
            target=self._drain, daemon=True)
        self._thread.start()

    def write(self, line):
        self.queue.put(line)   # non-blocking for workers

    def _drain(self):
        while True:
            line = self.queue.get()
            if line is None:
                return
            for stream in self.streams:
                stream.write(line)
                stream.flush()

🛡️ Workers Never Block

Worker threads enqueue and return immediately — no lock is held across I/O.

🔧 One I/O Thread

A single drain thread performs all writes, so a stall affects only that thread.

📢 Backpressure Warnings

The writer warns on stderr when the queue backs up, so a stall is never silent.

Aspect❌ Before✅ After
Lock held across writeYesNo
Buffer headroom~20 records (64 KB pipe)20,000 records (queue)
Stall visibilitySilent freezeExplicit warning
Consumer stall impactWhole pool freezesOne drain thread blocks