GitHub Actions skipped our production deploy, and no check went red

GitHub Actions skipped our blog's production deploy four times in a row on October 6, and no check in the workflow went red. A change we merged sat unshipped for five hours, until we started a fifth attempt by hand.

The change merged at 03:14 UTC. At 04:37, 83 minutes later, we checked the live blog. All three of its published posts still sent a header the change removes: a line of metadata the server sends with each page.

Our build job runs on self-hosted runners: machines we run ourselves, on our home internet line, instead of GitHub's. Each time, the build job hit its time limit while uploading the finished build, a 4.4 MB archive, for the deploy job. GitHub cancelled the build job. It showed as cancelled, not as a failure. Our production deploy job needs the build job to succeed first. With the build job cancelled, GitHub Actions skipped the deploy. A skipped job is not a failure either. GitHub's docs say a job that is skipped will report its status as "Success", and that it "will not prevent a pull request from merging, even if it is a required check." Nothing in the workflow flagged it.

One job at the end, built to turn a skipped deploy red

This applies if your deploy is its own GitHub Actions job that needs: your build job. When the build job fails or is cancelled, GitHub skips the deploy, unless the deploy job's if: says otherwise.

We added one job to the end of our deploy workflow. It needs both the build job, which we call verify, and the production deploy job. Here is the whole job except its script:

  production-deploy-alert:
    name: Report a production deploy that did not complete
    runs-on: light
    timeout-minutes: 5
    needs: [verify, deploy-production]
    if: always() && github.ref == 'refs/heads/main' && (github.event_name == 'push' || (github.event_name == 'workflow_dispatch' && inputs.environment == 'production'))
    concurrency:
      group: startupbros-funnels-production-deploy-alert
      queue: max
      cancel-in-progress: false
    permissions:
      actions: read
      issues: write
    steps:
      - name: Open, update or close the production deploy issue
        uses: actions/github-script@v9
        env:
          VERIFY_RESULT: ${{ needs.verify.result }}
          DEPLOY_RESULT: ${{ needs.deploy-production.result }}

light is the label of our own runners; use yours. The step runs actions/github-script, GitHub's action for running a script that uses the GitHub API from a workflow. GitHub hands the step the results of both jobs it needs, as VERIFY_RESULT and DEPLOY_RESULT. The script runs to 103 lines, and we have not published it.

The rest of the condition, after always(), limits the job to a push to main and to a production deploy started by hand.

Without a status check function such as always() in its if:, GitHub would skip this job along with the deploy. A job that needs a skipped job is skipped too, whatever status the skipped job reports. GitHub's page on needs says to use always() for a job that should run even if a job it depends on did not succeed. Its page on status check functions recommends !cancelled() instead, for a job or step that should run regardless of success or failure. We used always(). We have not tested whether !cancelled() also runs after a build that hit its time limit.

GitHub warns against always() on a task that could fail critically, such as getting the code, because the workflow may hang until it times out. Ours checks out no code and stops after five minutes.

What the step does, and what it won't:

  • The deploy did not succeed, and a newer push to main exists: it reports nothing. The newer run carries the change and reports for itself.
  • The deploy did not succeed, and no newer push exists: it opens an issue, or comments on the one already open, and fails the run.
  • The deploy succeeded: it closes the open issue, but only when this run is at least as new as the latest failure recorded on the issue. An older deploy that finishes late cannot close it.
  • Access: no checkout and no secret. Its GITHUB_TOKEN, the token GitHub gives the job, gets only the two permissions above. Once a job lists any permission, GitHub sets the ones it leaves out to none.
  • One at a time: alert jobs from different runs share one concurrency group, so, by design, one run does not read the issue while another is writing it. By default, GitHub lets one job wait in a group and cancels it when a newer one arrives. queue: max lets up to 100 wait, so a newer alert does not cancel an older one.

Rather have your agent write it? Paste this into Claude Code in your repo:

Add a last job to my deploy workflow. It needs my build job and my production deploy job, runs with if: always() on pushes to main, stops after five minutes, and checks out no code. Give it only actions: read and issues: write, and a concurrency group of its own with queue: max. If the deploy job did not succeed and no newer push to main exists, it opens an issue, or comments on the open one, and fails. If the deploy succeeded, it closes that issue, but only when this run is at least as new as the last failure the issue records. Show me the diff. Don't push.

By our reading of the workflow, it can raise a false alarm in one case we know of. A production deploy of main started by hand can cancel a push run's build job. That push run then finds no newer push and opens an issue. The deploy started by hand closes the issue when it succeeds, because its run is newer.

The job has run twice so far. Neither run was the case it exists for. The push that merged it, at 00:35 UTC on October 7, had a build that failed outright. That run was red already. After that failed build, the job opened an issue at 00:51. The next push to main deployed. The job then closed the issue at 07:05 UTC. Each run of the job took 8 seconds. No cancelled build has reached it yet.

Before you raise a time limit, read the step that hit it

We tried more time. We raised the build job's limit from 15 minutes to 60. The deploy run for that change became the fourth attempt. Its upload got as far as 97.3% of the archive. Three times, it fell back to 3.0% and started climbing again. It stood at 46.4% when GitHub cancelled the job at the one-hour mark.

The sign was in the step's log: the sent count went backwards. Here is the first fall-back, two lines a second apart:

2026-10-06T07:17:39.9951323Z Sent 4259840 of 4378496 (97.3%), 0.0 MBs/sec
2026-10-06T07:17:40.9955459Z Sent 131072 of 4378496 (3.0%), 0.0 MBs/sec

The upload looked like it was starting over, not inching toward the end. If the sent count in a step that hit its limit goes backwards, a longer limit may only move the cancellation later. Ours moved it from 15 minutes to an hour.

If your jobs run on your own machines, as ours do, measure packet loss and upload speed from the machine they run on before you raise the limit. Our early clean readings came from a computer that was not on the runners' line. They measured the wrong network.

Our best reading: a home internet line that was dropping packets

The build job hands the finished build to the deploy job through GitHub's cache, where one job saves files and a later job restores them, using actions/cache.

Before the change merged, a CI job on its pull request lost its self-hosted runner at 01:39:50 UTC, with GitHub's message "The self-hosted runner lost communication with the server." We measured loss on the home line twice on October 6, each time before an attempt. Between 02:01 and 02:20 UTC, before the first attempt, our router lost 48% of 1,300-byte test packets sent to the first router at our internet provider. At 06:30 UTC, before the fourth attempt, one runner host lost 42.5% of 1,300-byte test packets. Its test uploads ran at about 4.2 KB/s.

At 08:00 UTC, under four minutes after GitHub cancelled the 60-minute attempt at 07:56, we started a fifth attempt by hand. It deployed the same commit of main as the fourth attempt. That commit contained the change. The fifth attempt ran the same hand-off in 2 minutes 10 seconds, and the run as a whole shipped the change at 08:17. We have no reading of the line from during that attempt. We put the stalls down to the line, not the job. That rests on the two readings and the fast fifth attempt, not on a reading taken during any attempt.

We have since added an opt-in switch that moves the build and deploy jobs to a hosted runner provider's machines, off the home line. We merged it at 11:20 UTC on October 6, after the fifth attempt had shipped the change. None of the five attempts ran on it. It is off by default.

What this doesn't show

  • The alert reads a job's result, not the site. It reports that the production deploy job did not succeed. It never reads what production serves. If your deploy job can succeed without your change reaching users, this job will not tell you.
  • Seen live after one failed build, not after a cancelled one. Its one live failure came after a build that failed outright, in a run that was red anyway. Our tests run the job's script against a fake GitHub API, with a cancelled build among the cases, and check that it opens an issue and fails. Another test checks that its condition starts with always(). None of them runs the job on GitHub. GitHub has not yet run it after a real cancelled build. If you copy the job, the cancelled case is the one to test.
  • We have not fixed the line. As of October 7, the issue tracking the loss is still open. The opt-in switch goes around the line. Nothing in this post shows it carrying a deploy. The alert makes a skipped deploy visible. It does not ship the change.
  • Two readings, one day. The loss figures are two readings on October 6, one from our router and one from a runner host. They show the line was losing packets at those two times. They do not show how often it does, or what it did during each attempt. Read the cause in this post as our best reading, not as a measurement.

The blog in this story is House Notes, where we write up what our coding agents ran and measured.

Comments

No comments yet.