Fixing OpenAI API Timeouts in Laravel Jobs

October 1, 2026 · 4 views
Fixing OpenAI API Timeouts in Laravel Jobs

A Laravel job that calls the OpenAI API will eventually time out and get marked as failed, even though the request would have finished fine given another 20-30 seconds. This is one of the most common issues I run into when a project bolts an LLM feature onto an existing Laravel queue setup, and it's almost never actually a code bug. It's a mismatch between how long an OpenAI completion can realistically take and how the queue worker, the job class, and the HTTP client are each configured to wait for it.

I hit this exact thing on a client project a few months back: a "summarize this document" feature that worked fine in testing with short inputs, then started throwing MaxAttemptsExceededException in production the moment someone uploaded a 15-page PDF. The job wasn't broken. Three different timeout settings just disagreed with each other, and the shortest one won.

Why Laravel Queue Jobs Time Out on OpenAI Calls

When you dispatch a job that calls OpenAI::chat()->create() or hits the API directly through Guzzle, there are at least three separate clocks running, and Laravel doesn't reconcile them for you:

  1. The queue worker's own --timeout flag (defaults to 60 seconds), which sends a SIGALRM to kill the worker process if a single job runs longer than that.
  2. The job class's $timeout property, which overrides the worker flag per-job if you set it.
  3. The HTTP client's own timeout — Guzzle's timeout option, or whatever the OpenAI SDK sets by default, which is a completely separate clock that has nothing to do with Laravel's queue system at all.

A GPT-4-class model generating a long response, or a request competing with OpenAI's own rate limiting and retries under load, can easily take 30-90 seconds. If your worker timeout is still sitting at the Laravel default of 60 seconds and the job itself doesn't override it, the worker kills the process mid-request, Laravel marks the job as failed, and — depending on your retry_after setting — a second worker can pick the same job back up before the first attempt even finished dying. That's how you end up paying for the same OpenAI call twice.

The Settings That Actually Need to Agree

Here's a job class configured the way I'd actually set one up for an OpenAI call, with all three clocks pointed in the same direction:

class SummarizeDocument implements ShouldQueue
{
    use Dispatchable, InteractsWithQueue, Queueable, SerializesModels;

    public $timeout = 120;      // worker kills the job after this many seconds
    public $tries = 3;
    public $backoff = [10, 30, 60]; // seconds between retries

    public function handle(): void
    {
        $response = Http::timeout(90)        // Guzzle-level HTTP timeout
            ->connectTimeout(10)
            ->withToken(config('services.openai.key'))
            ->post('https://api.openai.com/v1/chat/completions', [
                'model' => 'gpt-4o',
                'messages' => $this->buildMessages(),
            ]);

        if ($response->failed()) {
            throw new OpenAiRequestFailed($response->body());
        }

        $this->document->update(['summary' => $response->json('choices.0.message.content')]);
    }
}

The rule I follow: the job's $timeout should always be longer than the HTTP client's timeout, never equal to it. Here the job gets 120 seconds, the HTTP call gets 90 — that gives Laravel 30 seconds of slack to catch a clean RequestException from Guzzle and fail the job gracefully, instead of the worker process getting killed mid-request by its own SIGALRM with no exception handling possible at all.

retry_after Is the Setting Most People Never Touch

Most tutorials tell you to just bump --timeout and move on. I'd skip that advice — on its own it doesn't fix anything, because retry_after in config/queue.php is a separate, independent clock that controls when a job becomes eligible to be picked up again by another worker, regardless of whether the first worker is still running it.

'connections' => [
    'database' => [
        'driver' => 'database',
        'table' => 'jobs',
        'queue' => 'default',
        'retry_after' => 150, // must exceed job $timeout, or you'll double-process
    ],
],

If retry_after is shorter than your job's $timeout, you get a genuinely nasty bug: a slow-but-successful OpenAI call is still running on worker A, retry_after expires, worker B grabs the "stuck" job and runs it again, and now you've billed two API calls and possibly written the summary twice depending on how idempotent your handle() method is. I've seen this specific failure mode cost a client real OpenAI spend before anyone noticed — set retry_after to at least $timeout + 30 and this class of bug goes away entirely.

Streaming Changes the Math Completely

A non-streamed chat/completions call blocks until OpenAI has generated the entire response, which is exactly the scenario above. If the job supports it, switching to a streamed response and writing tokens to the database (or broadcasting them over a websocket) as they arrive means the HTTP connection stays open and actively receiving data the whole time, rather than sitting idle waiting for one big payload — Guzzle's timeout applies to idle gaps between bytes in a stream, not total request duration, so a streamed job is far less likely to hit a false-positive timeout on a genuinely slow-but-healthy generation.

That said, streaming inside a queued job is awkward — you lose the simple "handle() returns, job is done" model and need somewhere to persist partial output if the worker itself dies mid-stream. For a background summarization job where nobody's watching it live, I'd still take the long non-streamed call with correctly-tuned timeouts over the added complexity of streaming into a job nobody's rendering in real time. Save streaming for requests actually driven by a user sitting in front of a browser.

A Pattern Worth Copying: Catch, Don't Let the Worker Kill It

The cleanest failure mode is one your own code controls, not one the OS controls via SIGALRM. Wrap the HTTP call so a real timeout becomes a handled exception your job's failed() method or Laravel's automatic retry logic can act on:

try {
    $response = Http::timeout(90)->post(/* ... */);
} catch (ConnectionException $e) {
    // Guzzle's own timeout fired — a clean, catchable exception
    $this->release(30); // put it back on the queue, try again shortly
    return;
}

A connection-level ConnectionException from Guzzle is something you can log, retry, or alert on. A worker getting killed by its own SIGALRM mid-request is not — you just get a vague failed-job entry with no real stack trace pointing at the actual cause.

Frequently Asked Questions

Does Laravel Horizon handle this differently than the database or Redis queue driver? Horizon still respects the same $timeout property and worker-level timeout settings — it's a dashboard and process supervisor on top of the same queue system, not a different timeout model. The retry_after gotcha applies the same way regardless of which driver you're running underneath it.

Should I just set a huge timeout like 300 seconds and stop worrying about it? You can, but it's a blunt fix. A genuinely hung request (network issue, OpenAI outage) now ties up a worker process for five minutes instead of failing fast, which matters a lot more than it sounds once you're running more than one or two workers. I'd rather tune the three settings to actually agree than mask the problem with a bigger number.

What's a reasonable OpenAI request timeout for a chat completion job? For GPT-4-class models on moderate-length inputs, 60-90 seconds at the HTTP client level is usually enough. If you're regularly summarizing long documents or generating long outputs, 120 seconds is more realistic — just make sure the job's $timeout and retry_after scale up to match.

Key Takeaways

Three clocks have to agree for a Laravel job calling OpenAI to fail cleanly instead of silently double-processing: the worker's --timeout, the job's $timeout, and the HTTP client's own timeout, with retry_after set comfortably above all of them. Get those four numbers pointed in the same direction, catch Guzzle's ConnectionException instead of letting the worker's SIGALRM do the killing, and a slow-but-healthy OpenAI call stops looking like a failure. If you only fix one thing from this post, fix retry_after — it's the setting responsible for the ugliest version of this bug, the silent double-charge.

#laravel #queues #php #openai-api #ai-integration #timeouts

0 Comments

No comments yet — be the first to share your thoughts.

Leave a comment

Never published.