What is it?
This case study covers how KineticSkunk designed, sized, and operated a custom GitLab runner fleet that replaced overloaded shared runners with specialised infrastructure tuned for workload isolation, build performance, and operational visibility.
The client outgrew default runner pools as queue times lengthened, flaky jobs multiplied, and security teams required isolation for artefacts on compliance-sensitive pipelines. FinOps wanted visibility into runner spend, not another opaque infrastructure bill line.
Use this approach when shared runners create queue contention, regulated workloads need isolation, build performance requires larger or specialised machines, and finance wants runner cost visibility alongside throughput.
Why it matters
Risks
- Shared runner pools create contention that blocks teams during peak build windows without clear prioritisation.
- Running compliance-sensitive workloads on shared infrastructure risks artefact and secret exposure across tenants.
- Unpatched runner hosts accumulate drift and vulnerabilities silently when they receive less attention than application servers.
Costs
- Queue contention on shared runners wastes developer waiting time that directly reduces feature delivery throughput.
- Oversized runners allocated for peak load sit idle most of the day, inflating infrastructure cost without proportional value.
- Transient runner failures that page operators unnecessarily consume incident response capacity.
Operational impact
- Without explicit job tagging, new pipelines land on default runners and compete with unrelated workloads for capacity.
- Manual runner recovery extends outages for builds and tests when self-healing automation is absent.
- Knowledge of runner fleet operations concentrates in one team member, creating a single point of failure for the entire CI/CD system.
Strategic impact
- Teams that control their runner infrastructure ship faster because queue times and build performance are predictable.
- Workload isolation on custom runners strengthens the compliance story for regulated industries.
- FinOps visibility into runner cost enables data-driven decisions about fleet sizing rather than guesswork.
How KineticSkunk completed the custom runner fleet engagement
Runner bottlenecks and isolation needs
- KineticSkunk assessed existing queue times, flake rates, and workload profiles to identify where shared runners were failing.
- Security and compliance requirements were mapped to determine which workloads needed dedicated, isolated infrastructure.
- FinOps data was gathered to establish baseline runner cost and identify oversized or underutilised capacity.
Fleet design and job routing
- Custom runner fleets were designed around workload profiles: build-heavy, test-heavy, compliance-isolated, and general-purpose.
- Explicit job tagging ensured predictable routing so new pipelines did not contend with unrelated workloads.
- Fleet sizing used real queue time data and workload patterns rather than theoretical peak capacity guessing.
Self-healing and operational automation
- Self-healing automation recovered runners from transient failures without paging operators for routine restarts.
- Patching and health checks received the same cadence and rigour as application server maintenance.
- Monitoring dashboards tied queue time, utilisation, and cost together so engineering and finance shared a single view.
Outcomes and ongoing operations
- Queue times dropped to predictable levels because workloads no longer competed for shared capacity.
- Pager noise reduced as self-healing handled transient failures that previously required manual intervention.
- Finance gained cost dashboards that showed runner utilisation beside throughput, enabling informed sizing decisions.
Common mistakes
Adding custom runners without explicit job tagging so all jobs land on them by default
Consequence: The new fleet inherits the same contention problem because unrelated workloads compete for the same capacity.
Avoidance: Tag jobs explicitly at creation so routing stays predictable when you add specialised fleets.
Sizing runners for theoretical peak load without examining actual queue time data
Consequence: Oversized runners sit idle most of the day, inflating cost without proportional throughput improvement.
Avoidance: Size fleets using real queue time and utilisation data, with auto-scaling for genuine burst demand.
Treating runner hosts as set-and-forget infrastructure that does not need patching
Consequence: Unpatched runners accumulate vulnerabilities and drift that compromise build artefact integrity and security.
Avoidance: Automate patching and health checks for runner hosts with the same rigour as application servers.
Best practices
- Tag jobs explicitly so routing stays predictable when you add specialised runner fleets.
- Automate patching and health checks for runner hosts with the same rigour as application servers.
- Keep cost dashboards beside queue time so finance sees the tradeoffs engineering optimises.
- Design self-healing automation so transient runner failures recover without operator intervention.
- Isolate compliance-sensitive workloads on dedicated runners to satisfy security and audit requirements.
- Size fleets using real queue time and utilisation data rather than theoretical peak capacity.
Tools and processes
- GitLab Runner with custom executors and fleet auto-scaling
- Explicit job tagging and runner tag matching for workload routing
- Self-healing automation with health checks and automatic restart
- FinOps dashboards combining queue time, utilisation, and cost metrics
- Patching automation aligned with application server maintenance cadence
How to get started
- Audit current queue times, flake rates, and workload profiles to identify shared runner limitations.
- Map compliance and isolation requirements to determine which workloads need dedicated runners.
- Design runner fleets around workload profiles with explicit job tagging for predictable routing.
- Implement self-healing automation and health checks to reduce operator burden.
- Build cost and queue time dashboards that give engineering and finance a shared operational view.
- Establish a patching cadence for runner hosts that matches application server maintenance standards.
If the immediate pain is queue contention, start with fleet sizing and job tagging. If the blocker is compliance isolation, start with dedicated runners for regulated workloads. Both paths converge on a well-operated custom fleet.
How KineticSkunk helps
KineticSkunk helps teams design, size, and operate custom GitLab runner fleets that deliver predictable build performance, workload isolation, and FinOps visibility.
The client achieved predictable queue times, reduced pager noise through self-healing automation, and gained cost visibility that enabled informed fleet sizing decisions.
For help designing runner fleets on GitLab, contact us or explore more case studies.
Frequently asked questions
When queue contention blocks delivery, compliance requires workload isolation, or build workloads need specialised hardware that shared pools cannot provide.
Explicit tags on jobs and runners ensure each pipeline step routes to the correct fleet, preventing unrelated workloads from competing for the same capacity.
Self-healing automation detects transient failures and restarts runner processes without paging operators, reducing noise and improving fleet availability.
Use real queue time and utilisation data rather than theoretical peak estimates. Add auto-scaling for burst demand and monitor cost alongside throughput.



