Four different problems hide behind one phrase
Bottleneck detection means four different things depending on who is asking. An operations leader means the order that sits in a queue for two days. An engineering lead means the API endpoint that takes four seconds under load. An engineering manager means the pull request that waits three days for a first review. A data lead means the nightly pipeline that finishes at 11am. The tools for each are different, and buying the wrong category is the most expensive mistake on this list.
I build and run production AI agents for operational teams, so my bias is towards the first kind. I have declared it, and I have tried to be fair to the others. No vendor on this page has paid to be here, and every price is a typical published range that you should verify with the vendor before budgeting.
1. Celonis
Celonis is the market leader in process mining. It ingests event logs from your ERP and CRM, reconstructs how work actually flows, and shows you where cases queue, loop back or get abandoned. When it works, it is genuinely revealing. Teams discover that the approved process and the real process diverged years ago.
It wins in large enterprises running SAP or Oracle with high case volumes, where the event data already exists and a one percent improvement pays for the licence. It is the wrong buy for a mid-size company whose process lives in email threads, spreadsheets and people's heads, because there is no event log to mine. Engagements are typically six figures annually once implementation is included, on custom quotes.
2. UiPath Process Mining
UiPath bought ProcessGold years ago and folded process mining into its automation platform. The pitch is coherent. Mine the process, find the bottleneck, then automate it with the RPA tooling you already own. If you are already a UiPath shop, the incremental cost and learning curve are lower than a standalone Celonis deployment.
It wins when process mining feeds directly into an existing automation programme. It is the wrong buy as a standalone mining tool, because you are paying for platform gravity you will not use. Pricing is bundled into UiPath platform quotes, so get the mining line item priced separately before you sign.
3. Datadog
If your bottleneck is software, Datadog is the default answer for a reason. APM traces show you exactly which function, query or downstream call is eating the latency budget, with flame graphs that make the slow path visible in minutes. Database monitoring and profiling close the loop.
It wins for engineering teams running production services who need answers under incident pressure. It is the wrong buy for business-process bottlenecks, which never appear in a trace. The other caution is cost creep. Published pricing, billed annually, starts around 15 to 23 dollars per host per month for infrastructure and roughly 31 to 40 dollars per host per month for APM, but real bills grow with hosts, custom metrics and log volume, so model a year of growth before committing.
4. Grafana with Prometheus and Tempo
The open-source observability stack does most of what Datadog does if you have engineers willing to run it. Prometheus for metrics, Tempo or Jaeger for traces, Grafana on top for dashboards and alerting. Grafana Cloud offers a generous free tier and usage-based paid plans if you would rather not self-host everything.
It wins for engineering-led teams who want control and predictable cost. It is the wrong buy if nobody owns it, because a neglected self-hosted stack quietly becomes its own bottleneck. Budget the engineering time honestly. Free software is not free operation.
5. Data-stack profiling
Sometimes the bottleneck is the data platform itself. Reports arrive late because the nightly dbt run takes six hours, or one unindexed join in the warehouse burns the whole window. The tools here are mostly ones you already have. Snowflake query profiles, BigQuery execution plans, dbt run timings, and observability layers such as Monte Carlo or Elementary if you want alerting on top.
This wins whenever the complaint is that the numbers are late rather than wrong. It is less a purchase than a discipline, which is exactly why it gets skipped. Before buying anything new, spend a day reading the query profiles you already pay for.
6. Ops analytics dashboards
Most operational systems already record timestamps. Order created, order approved, order shipped. A cycle-time dashboard in Power BI or Metabase built on those timestamps will find your biggest process bottleneck for the cost of a few analyst days. I compare the AI-flavoured versions of these tools in my guide to AI analytics tools for mid-size companies.
This wins as the first serious step beyond gut feel, and for many mid-size companies it is enough. The limitation is that dashboards show lagging averages and only help if someone looks at them. The queue that blows up on Tuesday is invisible until the Monday review.
7. Spreadsheets plus scripts
The honest baseline, and I mean that without irony. Export the timestamps, load them into a spreadsheet or a short Python script, and compute time-in-stage for every case over the last quarter. Sort descending. The bottleneck is usually staring at you within an afternoon, and the exercise costs nothing but focus.
It wins for a first diagnosis, for small teams, and for validating whatever a vendor demo just showed you. It is the wrong tool as a permanent answer, because the script rots the moment the process changes and nobody re-runs it. Treat it as reconnaissance, not infrastructure.
8. A production AI agent watching the process
This is the category I work in, so weigh my view accordingly. Instead of mining history or reviewing dashboards weekly, an agent connects to the same systems your team uses, recomputes stage times continuously, flags a queue the moment it starts growing abnormally, and drafts the chase or the escalation itself. The detection and the response collapse into one loop.
It wins for mid-size operational teams that have no data function and need the watching done for them, not more charts. It is the wrong buy if you have not first done the spreadsheet exercise, because you will not know what normal looks like, and it fails without real system access and a clear owner. I have written about why the model is never the hard part of these systems, and about what a production system owes you that a feature never will. Both apply doubly here.
When the bottleneck is software delivery
The fourth meaning is the one engineering managers search for. The code is written, but the pull request waits for a first review, the build queues behind others, and the release sits until the next window. None of that shows in an APM trace, because nothing is slow. It is waiting. The tools below split the time from first commit to production into stages, so you can see which stage holds the work.
The same caution applies as everywhere else on this page. The data is already in your repository and your build system, and a free first pass tells you which stage is the problem before any vendor does. I have written about why developers are rarely the bottleneck in software delivery and about measuring queue time rather than touch time.
9. The analytics in your build tool
Start here, because you may already pay for it. GitHub Actions has shown queue time, run time and failure rates on every GitHub Cloud plan since March 2025. GitLab Value Stream Analytics shows the median time work spends in each stage, with a default view on the free tier at project level, while custom value streams need Premium and the DORA dashboards need Ultimate. CircleCI Insights shows build duration percentiles and success rates. If you already run Datadog, CI Visibility adds pipeline queue time and the critical path for a published 8 dollars per active committer per month, billed annually.
This wins when the complaint is slow or flaky builds, and it costs little or nothing to look. It is the wrong place to stop if the wait is human, because a build dashboard cannot see a pull request nobody has picked up for review.
10. LinearB
LinearB splits cycle time into four stages. Coding runs from first commit to pull request, pickup from pull request to first review, review from first review to merge, and deploy from merge to production. The pickup figure is usually the revealing one, because it is pure waiting. It also tracks DORA metrics and pull request size.
It wins for engineering organisations large enough to have several teams and a recurring argument about where delivery time goes. It is the wrong buy for a small team, since the entry plan is published at 29 dollars per contributor per month billed annually and is built for teams of 50 or more developers, on GitHub Cloud only. The 45-day trial is long enough to answer the question before you commit.
11. Swarmia
Swarmia shows the same split from the team's side, with time to first review, review time and time to merge, alongside CI visibility for the slowest pipelines, DORA metrics and working agreements a team sets for itself, such as a target time to review.
It wins for teams that want to change habits, not only report on them, and it is free for companies with up to nine developers, which makes it the cheapest way to try a dedicated engineering analytics tool. Paid plans are published at 45 dollars per developer per month billed annually, so price the whole team before the trial ends. It is the wrong buy if nobody will act on what it shows, because its value is in the habits it changes.
12. Jellyfish and Atlassian (DX)
These two sit at the enterprise end. Gartner published its first Magic Quadrant for developer productivity insight platforms in May 2026, and the vendors' own announcements say it named Atlassian (DX), Jellyfish, LinearB and Opsera as leaders. Jellyfish splits delivery into refinement, work, review and deployment using Jira and Git data, and adds a view of where engineering investment goes. DX, which Atlassian bought in November 2025, combines developer surveys with system data, so it catches friction that never shows up in a timestamp.
They win for large engineering organisations that need delivery data to line up with finance and planning. Neither publishes pricing, so expect a sales process. They are the wrong buy when the question is one team's review queue, which a free export answers in an afternoon.
Two names still appear in older lists. Pluralsight Flow is being discontinued by its new owner, with no renewals past 30 June 2026 and end of life on 31 December 2027. Code Climate Velocity has been replaced by a custom platform sold with services. Neither is worth starting with now.
Comparison at a glance
| Tool | Bottleneck type | Typical cost | Time to first insight | Wrong buy when |
|---|---|---|---|---|
| Celonis | Business process | Six figures per year, custom quote | Months | No clean event logs, mid-size budget |
| UiPath Process Mining | Business process | Bundled platform quote | Months | Not already on UiPath |
| Datadog | Software | Roughly 15 to 40 dollars per host per month, grows fast | Days | Bottleneck is a process, not code |
| Grafana stack | Software | Free tier, then usage-based, plus engineer time | Days to weeks | Nobody owns the stack |
| Data-stack profiling | Data pipeline | Mostly included in what you pay already | Hours | Rarely, it is near free |
| Ops dashboards | Business process | Analyst days plus BI seats | Days | Nobody will review them |
| Spreadsheets plus scripts | Any, once | Nearly nothing | An afternoon | Used as a permanent system |
| AI agent on the process | Business process, continuous | Build plus run cost, scoped per process | Weeks | No baseline, no system access, no owner |
| Build tool analytics | Build and release pipeline | Included in GitHub, GitLab and CircleCI plans, Datadog CI from 8 dollars per committer per month | Hours | The wait is a person, not a build |
| LinearB | Software delivery | From 29 dollars per contributor per month, built for 50 or more developers | Days | Small team, one-off question |
| Swarmia | Software delivery | Free up to 9 developers, then 45 dollars per developer per month | Days | Nobody will change review habits |
| Jellyfish, Atlassian (DX) | Software delivery at scale | Custom quote | Weeks | One team's review queue |
How to pick by company size
Under roughly 50 people, do not buy anything. Run the spreadsheet exercise, fix the worst queue, repeat quarterly. Your bottlenecks are usually visible from the founder's chair once the timestamps are in front of you. For an engineering team the same exercise uses pull request timestamps, and the analytics in your build tool cover the rest.
Between roughly 50 and 500 people, build the ops dashboard first and profile your data stack the same week. If the same bottleneck keeps reappearing and someone is spending hours a week chasing it, that is the point where a continuously watching agent earns its keep, because the cost of the human loop now exceeds the cost of the system. On the engineering side this is the range where a per-seat delivery analytics tool starts to pay for itself, provided someone owns acting on what it shows.
Above that, process mining starts to make sense, provided your core systems generate usable event logs and you have an owner for the programme. Pair it with proper APM for the software side and an engineering analytics platform for delivery. At every size, the failure mode is the same one I keep seeing in AI projects generally, which is buying the tool before doing the unglamorous groundwork. If that pattern sounds familiar, read why pilots look great and then die before production before you sign anything.
More field notes like this live on the insights page.