The pipeline is costing you more than your slowest engineer
Most engineering leaders I talk to have already looked hard at their developers. Sprint velocity, PR cycle time, review turnaround. What they have not looked at is how many engineer-hours per week evaporate while a build sits in a queue, reruns a test that fails one time in five for no reproducible reason, or crawls through a stage that could have run in parallel an hour ago.
In a mid-sized distribution business I worked with, the engineering team had roughly forty developers pushing to a shared monorepo. Their measured median time from merge to deployable artefact was around ninety minutes. About twenty minutes of that was actual compilation and test execution. The other seventy minutes was waiting. Waiting for a runner to become available, waiting for a flaky integration test to be retried, waiting for a security scan that ran after unit tests instead of alongside them. Nobody had ever added those numbers up and put them on a slide. When we did, the room went quiet.
That is the pattern I see repeatedly across finance, manufacturing and distribution engineering teams. The build pipeline is treated as infrastructure, not as a delivery constraint. It gets attention when it is completely broken and almost none otherwise.
How to measure wait time versus build time before you touch anything
Before you fix anything, measure the right thing. Most CI platforms expose enough event data to reconstruct two numbers for every pipeline run, time in queue and time executing. If your platform does not surface these directly, you can derive them from job-created timestamps versus job-started timestamps. That gap is your queue time. Everything from job-started to job-finished is your execution time. Do this across a rolling thirty-day window and bucket the results by branch type, by time of day and by pipeline stage.
What you are looking for is the ratio. In a healthy pipeline, queue time is a small fraction of execution time, somewhere in the range of ten to twenty percent. When I see queue time running at fifty percent or more of total pipeline duration, that is a runner capacity problem and no amount of test optimisation will fix it. When queue time is low but execution time is high and variable, that points to flaky tests or poorly ordered stages.
A pattern I have seen in financial services engineering teams, particularly those running compliance and audit pipelines alongside feature pipelines, is that queue time spikes sharply between roughly nine in the morning and noon because every team pushes before the daily standup. The fix there is not more runners across the board. It is autoscaling with a warm pool sized to the morning spike, and staggering the compliance pipeline to run on a slight delay so it does not compete with feature pipelines for the same runner pool at the same moment.
Measure first. The order of fixes depends entirely on what the data shows.
Flaky tests are not a test problem they are a trust problem
A flaky test is one that produces different results on the same code without any change to the code. This sounds like a testing problem. It is actually a trust problem, and trust problems compound.
Here is what happens in practice. A test fails. The developer reruns the pipeline. It passes. They merge. Over time the team learns that a certain class of failures are noise. They start ignoring red builds. Then a real regression slips through because the signal-to-noise ratio has degraded to the point where nobody trusts the red state anymore. I have seen this pattern in manufacturing engineering teams where integration tests against a simulated PLC environment would fail intermittently due to timing issues in the test harness rather than anything wrong with the application code. The team had accumulated somewhere in the range of thirty to forty tests that were known to be flaky. They were retried automatically. The retry logic masked the problem for months until a genuine defect in a conveyor control routine also got masked and made it to staging.
The fix for flaky tests has a specific order. First, quarantine them. Move them to a separate suite that runs but does not gate the pipeline. This removes the noise from the main signal immediately. Second, classify them. Timing-dependent tests, tests with external dependencies that are not mocked, tests that share mutable state across runs. Each class has a different fix. Third, fix or delete them in order of how often they fire, not in order of how important the thing they test seems to be. A flaky test that fires three times a day costs more than a solid test covering something critical.
Do not retry flaky tests as a permanent strategy. Retries are a diagnostic tool. If a test needs a retry to pass, that test has a defect. Treat it as such.
Runner queues and why throwing more capacity at them usually does not work
When developers complain that the pipeline is slow, the first instinct in most organisations is to add more runners. Sometimes that is right. Often it is not, and adding runners without understanding the queue structure just moves the bottleneck somewhere else.
Runner queues form for two distinct reasons and they require different responses. The first is genuine throughput shortage, where more pipeline runs arrive per unit time than the runner pool can process. The second is structural inefficiency, where runners sit idle while jobs wait because of tagging, affinity rules or stage dependencies that prevent available runners from picking up waiting jobs.
In a distribution business I worked with, the team had a pool of runners that looked adequately sized on average utilisation metrics. Average utilisation was around forty percent. But peak utilisation during the pre-release window, which happened to coincide with end-of-quarter logistics reporting runs, hit one hundred percent and queues built up to forty-five minutes. The fix was not a larger static pool. It was autoscaling with a ceiling sized to the peak and a floor sized to the overnight baseline, combined with priority lanes so release-blocking pipelines could jump the queue ahead of scheduled background jobs.
The structural inefficiency version is more insidious. I have seen pipelines in financial services teams where a runner tagged for a specific compliance toolchain sat idle for twenty minutes while a job that needed exactly that runner waited in queue, because the tagging rules had been set up years earlier and nobody had revisited them as the toolchain evolved. The job could have run on three other runner types but the tag prevented it. Auditing your runner tags and affinity rules is unglamorous work. It is also frequently worth an hour or two of pipeline time per day.
Serial stages are usually the last thing to fix and the first thing people try
Parallelising pipeline stages is the most visible intervention and the one that gets proposed first in almost every conversation I have about slow pipelines. It is also usually the third thing to fix, not the first.
Here is why the order matters. If your pipeline has significant queue time, parallelising stages does not help because your jobs are waiting for runners, not waiting for each other. If your pipeline has significant flaky test noise, parallelising stages spreads that noise across more concurrent jobs and makes the signal harder to read. You fix queue time first, flaky tests second, and then you look at stage parallelisation.
When you do get to stage ordering, the analysis is straightforward. Map every stage, its duration and its dependencies. Draw the critical path. Any stage that is not on the critical path and is currently running serially is a candidate for parallelisation. Any stage that is on the critical path and has no hard dependency on a preceding stage is also a candidate.
A manufacturing engineering team I worked with had a pipeline that ran unit tests, then integration tests, then a static analysis pass, then a container build, then a smoke test against the built container. The static analysis had no dependency on the integration test results. Moving it to run in parallel with integration tests took roughly twelve minutes off the critical path. The container build had no dependency on the static analysis result either, only on the unit tests. Restructuring the dependency graph so the container build started as soon as unit tests passed, with integration tests and static analysis running in parallel alongside it, took another eight minutes off. Twenty minutes of improvement from dependency graph work, not from adding any resources.
The caveat is that parallel stages increase the complexity of your pipeline definition and your failure diagnosis. A serial pipeline that fails is easy to read. A parallel pipeline that fails requires you to look at multiple concurrent job logs. Make sure your observability is good enough to support the complexity before you add it.
The order to fix things and why it matters
Based on what I have seen across engineering teams in finance, manufacturing and distribution, the right order is this.
- Measure queue time versus execution time across a meaningful window. If queue time is above thirty percent of total pipeline duration, fix runner capacity or structural runner inefficiency before anything else.
- Quarantine and classify flaky tests. Do not retry them. Do not ignore them. Quarantine them so they stop poisoning the signal, then fix them in order of firing frequency.
- Audit runner tags and affinity rules. This is cheap and often yields significant queue time reduction without adding any capacity.
- Parallelise stages by mapping the critical path and eliminating unnecessary serial dependencies. Do this after queue time and flaky tests are under control.
- Set a pipeline duration budget and treat breaches as engineering incidents. Without a budget, pipelines drift slow over time as stages are added and nobody removes anything.
The reason order matters is that each fix changes the baseline for the next measurement. If you parallelise stages before fixing runner queues, you will see less improvement than the dependency graph analysis suggests because the jobs are still waiting for runners. If you add runners before quarantining flaky tests, you will process more retries faster and your costs will go up without your reliability going up.
Practical takeaway
The single most useful thing an engineering or operations leader can do this week is pull thirty days of pipeline run data and compute the ratio of queue time to execution time for every stage. That number tells you immediately whether you have a capacity problem, a test reliability problem or a stage ordering problem. Most teams have never looked at it.
A pipeline that takes ninety minutes but only does twenty minutes of real work is not a slow pipeline. It is a pipeline that is waiting for seventy minutes. Those are fixable in different ways and for different costs. Know which one you have before you spend anything.
The teams I have seen make the most improvement are the ones that treat pipeline duration as a first-class engineering metric with an owner, a budget and a review cadence. Not a background concern. Not something that gets attention only when a release is on fire. A constraint on delivery capacity that gets the same rigour as any other constraint in the system.
