Alexey Shurov.Insights
Tools

Top tools for bottleneck detection in 2026

Process mining, observability, engineering analytics, ops dashboards and AI agents all claim to find your bottlenecks. Here is what each one is actually for, and when it is the wrong buy.

Updated 25 September 2026 . 11 min read . Alexey Shurov

Four different problems hide behind one phrase

Bottleneck detection means four different things depending on who is asking. An operations leader means the order that sits in a queue for two days. An engineering lead means the API endpoint that takes four seconds under load. An engineering manager means the pull request that waits three days for a first review. A data lead means the nightly pipeline that finishes at 11am. The tools for each are different, and buying the wrong category is the most expensive mistake on this list.

I build and run production AI agents for operational teams, so my bias is towards the first kind. I have declared it, and I have tried to be fair to the others. No vendor on this page has paid to be here, and every price is a typical published range that you should verify with the vendor before budgeting.

1. Celonis

Celonis is the market leader in process mining. It ingests event logs from your ERP and CRM, reconstructs how work actually flows, and shows you where cases queue, loop back or get abandoned. When it works, it is genuinely revealing. Teams discover that the approved process and the real process diverged years ago.

It wins in large enterprises running SAP or Oracle with high case volumes, where the event data already exists and a one percent improvement pays for the licence. It is the wrong buy for a mid-size company whose process lives in email threads, spreadsheets and people's heads, because there is no event log to mine. Engagements are typically six figures annually once implementation is included, on custom quotes.

2. UiPath Process Mining

UiPath bought ProcessGold years ago and folded process mining into its automation platform. The pitch is coherent. Mine the process, find the bottleneck, then automate it with the RPA tooling you already own. If you are already a UiPath shop, the incremental cost and learning curve are lower than a standalone Celonis deployment.

It wins when process mining feeds directly into an existing automation programme. It is the wrong buy as a standalone mining tool, because you are paying for platform gravity you will not use. Pricing is bundled into UiPath platform quotes, so get the mining line item priced separately before you sign.

3. Datadog

If your bottleneck is software, Datadog is the default answer for a reason. APM traces show you exactly which function, query or downstream call is eating the latency budget, with flame graphs that make the slow path visible in minutes. Database monitoring and profiling close the loop.

It wins for engineering teams running production services who need answers under incident pressure. It is the wrong buy for business-process bottlenecks, which never appear in a trace. The other caution is cost creep. Published pricing, billed annually, starts around 15 to 23 dollars per host per month for infrastructure and roughly 31 to 40 dollars per host per month for APM, but real bills grow with hosts, custom metrics and log volume, so model a year of growth before committing.

4. Grafana with Prometheus and Tempo

The open-source observability stack does most of what Datadog does if you have engineers willing to run it. Prometheus for metrics, Tempo or Jaeger for traces, Grafana on top for dashboards and alerting. Grafana Cloud offers a generous free tier and usage-based paid plans if you would rather not self-host everything.

It wins for engineering-led teams who want control and predictable cost. It is the wrong buy if nobody owns it, because a neglected self-hosted stack quietly becomes its own bottleneck. Budget the engineering time honestly. Free software is not free operation.

5. Data-stack profiling

Sometimes the bottleneck is the data platform itself. Reports arrive late because the nightly dbt run takes six hours, or one unindexed join in the warehouse burns the whole window. The tools here are mostly ones you already have. Snowflake query profiles, BigQuery execution plans, dbt run timings, and observability layers such as Monte Carlo or Elementary if you want alerting on top.

This wins whenever the complaint is that the numbers are late rather than wrong. It is less a purchase than a discipline, which is exactly why it gets skipped. Before buying anything new, spend a day reading the query profiles you already pay for.

6. Ops analytics dashboards

Most operational systems already record timestamps. Order created, order approved, order shipped. A cycle-time dashboard in Power BI or Metabase built on those timestamps will find your biggest process bottleneck for the cost of a few analyst days. I compare the AI-flavoured versions of these tools in my guide to AI analytics tools for mid-size companies.

This wins as the first serious step beyond gut feel, and for many mid-size companies it is enough. The limitation is that dashboards show lagging averages and only help if someone looks at them. The queue that blows up on Tuesday is invisible until the Monday review.

7. Spreadsheets plus scripts

The honest baseline, and I mean that without irony. Export the timestamps, load them into a spreadsheet or a short Python script, and compute time-in-stage for every case over the last quarter. Sort descending. The bottleneck is usually staring at you within an afternoon, and the exercise costs nothing but focus.

It wins for a first diagnosis, for small teams, and for validating whatever a vendor demo just showed you. It is the wrong tool as a permanent answer, because the script rots the moment the process changes and nobody re-runs it. Treat it as reconnaissance, not infrastructure.

8. A production AI agent watching the process

This is the category I work in, so weigh my view accordingly. Instead of mining history or reviewing dashboards weekly, an agent connects to the same systems your team uses, recomputes stage times continuously, flags a queue the moment it starts growing abnormally, and drafts the chase or the escalation itself. The detection and the response collapse into one loop.

It wins for mid-size operational teams that have no data function and need the watching done for them, not more charts. It is the wrong buy if you have not first done the spreadsheet exercise, because you will not know what normal looks like, and it fails without real system access and a clear owner. I have written about why the model is never the hard part of these systems, and about what a production system owes you that a feature never will. Both apply doubly here.

When the bottleneck is software delivery

The fourth meaning is the one engineering managers search for. The code is written, but the pull request waits for a first review, the build queues behind others, and the release sits until the next window. None of that shows in an APM trace, because nothing is slow. It is waiting. The tools below split the time from first commit to production into stages, so you can see which stage holds the work.

The same caution applies as everywhere else on this page. The data is already in your repository and your build system, and a free first pass tells you which stage is the problem before any vendor does. I have written about why developers are rarely the bottleneck in software delivery and about measuring queue time rather than touch time.

9. The analytics in your build tool

Start here, because you may already pay for it. GitHub Actions has shown queue time, run time and failure rates on every GitHub Cloud plan since March 2025. GitLab Value Stream Analytics shows the median time work spends in each stage, with a default view on the free tier at project level, while custom value streams need Premium and the DORA dashboards need Ultimate. CircleCI Insights shows build duration percentiles and success rates. If you already run Datadog, CI Visibility adds pipeline queue time and the critical path for a published 8 dollars per active committer per month, billed annually.

This wins when the complaint is slow or flaky builds, and it costs little or nothing to look. It is the wrong place to stop if the wait is human, because a build dashboard cannot see a pull request nobody has picked up for review.

10. LinearB

LinearB splits cycle time into four stages. Coding runs from first commit to pull request, pickup from pull request to first review, review from first review to merge, and deploy from merge to production. The pickup figure is usually the revealing one, because it is pure waiting. It also tracks DORA metrics and pull request size.

It wins for engineering organisations large enough to have several teams and a recurring argument about where delivery time goes. It is the wrong buy for a small team, since the entry plan is published at 29 dollars per contributor per month billed annually and is built for teams of 50 or more developers, on GitHub Cloud only. The 45-day trial is long enough to answer the question before you commit.

11. Swarmia

Swarmia shows the same split from the team's side, with time to first review, review time and time to merge, alongside CI visibility for the slowest pipelines, DORA metrics and working agreements a team sets for itself, such as a target time to review.

It wins for teams that want to change habits, not only report on them, and it is free for companies with up to nine developers, which makes it the cheapest way to try a dedicated engineering analytics tool. Paid plans are published at 45 dollars per developer per month billed annually, so price the whole team before the trial ends. It is the wrong buy if nobody will act on what it shows, because its value is in the habits it changes.

12. Jellyfish and Atlassian (DX)

These two sit at the enterprise end. Gartner published its first Magic Quadrant for developer productivity insight platforms in May 2026, and the vendors' own announcements say it named Atlassian (DX), Jellyfish, LinearB and Opsera as leaders. Jellyfish splits delivery into refinement, work, review and deployment using Jira and Git data, and adds a view of where engineering investment goes. DX, which Atlassian bought in November 2025, combines developer surveys with system data, so it catches friction that never shows up in a timestamp.

They win for large engineering organisations that need delivery data to line up with finance and planning. Neither publishes pricing, so expect a sales process. They are the wrong buy when the question is one team's review queue, which a free export answers in an afternoon.

Two names still appear in older lists. Pluralsight Flow is being discontinued by its new owner, with no renewals past 30 June 2026 and end of life on 31 December 2027. Code Climate Velocity has been replaced by a custom platform sold with services. Neither is worth starting with now.

Comparison at a glance

ToolBottleneck typeTypical costTime to first insightWrong buy when
CelonisBusiness processSix figures per year, custom quoteMonthsNo clean event logs, mid-size budget
UiPath Process MiningBusiness processBundled platform quoteMonthsNot already on UiPath
DatadogSoftwareRoughly 15 to 40 dollars per host per month, grows fastDaysBottleneck is a process, not code
Grafana stackSoftwareFree tier, then usage-based, plus engineer timeDays to weeksNobody owns the stack
Data-stack profilingData pipelineMostly included in what you pay alreadyHoursRarely, it is near free
Ops dashboardsBusiness processAnalyst days plus BI seatsDaysNobody will review them
Spreadsheets plus scriptsAny, onceNearly nothingAn afternoonUsed as a permanent system
AI agent on the processBusiness process, continuousBuild plus run cost, scoped per processWeeksNo baseline, no system access, no owner
Build tool analyticsBuild and release pipelineIncluded in GitHub, GitLab and CircleCI plans, Datadog CI from 8 dollars per committer per monthHoursThe wait is a person, not a build
LinearBSoftware deliveryFrom 29 dollars per contributor per month, built for 50 or more developersDaysSmall team, one-off question
SwarmiaSoftware deliveryFree up to 9 developers, then 45 dollars per developer per monthDaysNobody will change review habits
Jellyfish, Atlassian (DX)Software delivery at scaleCustom quoteWeeksOne team's review queue

How to pick by company size

Under roughly 50 people, do not buy anything. Run the spreadsheet exercise, fix the worst queue, repeat quarterly. Your bottlenecks are usually visible from the founder's chair once the timestamps are in front of you. For an engineering team the same exercise uses pull request timestamps, and the analytics in your build tool cover the rest.

Between roughly 50 and 500 people, build the ops dashboard first and profile your data stack the same week. If the same bottleneck keeps reappearing and someone is spending hours a week chasing it, that is the point where a continuously watching agent earns its keep, because the cost of the human loop now exceeds the cost of the system. On the engineering side this is the range where a per-seat delivery analytics tool starts to pay for itself, provided someone owns acting on what it shows.

Above that, process mining starts to make sense, provided your core systems generate usable event logs and you have an owner for the programme. Pair it with proper APM for the software side and an engineering analytics platform for delivery. At every size, the failure mode is the same one I keep seeing in AI projects generally, which is buying the tool before doing the unglamorous groundwork. If that pattern sounds familiar, read why pilots look great and then die before production before you sign anything.

More field notes like this live on the insights page.

Common questions

What software helps identify pipeline bottlenecks?

It depends which pipeline. For a build and release pipeline, start with the analytics in your CI tool, GitHub Actions performance metrics, GitLab Value Stream Analytics or CircleCI Insights, which show queue time, run time and failure rate. For a data pipeline, read the query profiles and run timings in the warehouse and orchestration tools you already have. For a business process, a cycle-time dashboard on the timestamps your systems already record comes first, and process mining makes sense at enterprise scale.

What is the best software for finding engineering bottlenecks?

For delivery, engineering analytics platforms such as LinearB, Swarmia, Jellyfish and Atlassian (DX) split cycle time into coding, waiting for review, review and deploy, which is where delivery time usually hides. If the problem is slow code rather than slow delivery, APM such as Datadog or the Grafana stack is the right category instead. Before buying either, export when each pull request was opened, first reviewed and merged over the last quarter and compute the time in each stage, which answers the question for free.

How do you find the bottleneck in an SDLC workflow?

Split the time from first commit to production into stages and measure the waiting separately from the working. The stage where work waits longest is the constraint. Often it is the wait for a first review or the release process rather than the coding itself. Fix that stage, then measure again, because the constraint moves.

Can AI detect workflow bottlenecks and friction?

Within limits, yes. An agent connected to the systems a team already uses can recompute stage times continuously, flag a queue the moment it grows abnormally and spot chasers and repeated questions in shared inboxes. It needs system access, a baseline of what normal looks like and an owner who acts on what it flags. It will not tell you why a step exists, which is still a conversation with the people who run it.

How much do bottleneck detection tools cost?

From nothing to six figures a year. Spreadsheets, build tool analytics and Swarmia's free tier for up to nine developers cost little or nothing. Published per-seat pricing for engineering analytics starts at 29 dollars per contributor per month for LinearB, aimed at teams of 50 or more developers, and 45 dollars per developer per month for Swarmia, both billed annually. Datadog infrastructure monitoring starts at 15 dollars per host per month billed annually. Enterprise process mining is sold on custom quotes that commonly reach six figures a year once implementation is included. Check current prices with each vendor before budgeting.

Want the watching done for you

I build and run production AI agents that monitor operational processes and take the repetitive chasing off your team. Tell me where your work queues up.

More insightsshurco.ai