Scaling platform engineering through AI-powered delivery agents

How a leading professional services organisation accelerated platform onboarding, strengthened compliance and enabled self-serve CI/CD support through AI

Our client – a leading global professional services organisation – wanted to help development teams onboard services faster, resolve CI/CD issues more independently, and improve visibility of security compliance across a shared deployment platform. Working closely with the client team, Equal Experts explored how to make platform knowledge, delivery workflows and compliance controls more accessible through specialised AI agents, without changing the underlying operating model.

At the time, the organisation was running over 1,000 applications on a shared deployment platform. By design, application developers did not have direct access to the platform, and diagnosing issues often depended on expertise held within the specialist team. For example, if a deployment were to fail during a live feature release, the platform team would need to get involved, often spending up to 2 hours diagnosing and resolving the problem. With teams working across different time zones, this often slowed down delivery and created frustration on both sides.

Together, we introduced four specialised AI agents, embedded into the organisation’s existing tools. These agents supported developers and platform teams across the build, deploy and run phases, generating compliant service configurations, diagnosing and troubleshooting deployment failures, and giving security teams a unified view across all services.

Outcomes

100+

services onboarded per quarter (up from ~25)

80%

of pipeline failures resolved without specialist intervention, reducing average incident resolution time from ~2 hours to minutes

45 minutes

to onboard new services (down from 2 days)

About the client

The client is a leading professional services organisation delivering solutions through technology, consulting and design.

Industry
Consulting
Organisation size
$10bn revenue and over 30k employees
Location
Global
Length of project
Ongoing

Challenge

Onboarding, diagnostics and compliance: bottlenecked on platform specialists

When we began working with the client, diagnosing a pipeline failure required correlating logs across multiple systems, understanding which tools were involved and interpreting a wide range of error states. For application developers, access was the first barrier. By design (and in line with organisational policy), they could not directly access the platform. Beyond that, resolving issues depended on specialist knowledge that was not widely shared.

Even developers who had some familiarity with the platform often lacked the breadth of understanding needed to diagnose the full range of failure types. These included configuration errors, dependency issues, permission mismatches, environment inconsistencies, and tool-specific problems.

In addition, onboarding a service was a multi-day coordination effort to assemble a compliant build-and-deploy configuration.

Frequent changes to the tooling landscape – with major transitions taking place roughly every six months – compounded the issue, keeping diagnostic knowledge perpetually concentrated in a small number of people and out of reach for everyone else.

Solution

Embedding platform expertise through secure AI agents

The expertise needed to onboard services, diagnose failures and verify compliance already existed, but it was simply held by a handful of specialists and only available when they were directly involved. The goal was to make that knowledge available on demand, embedded in the tools developers already used, without putting those individuals in the loop and without weakening the platform’s access controls.

Working with the client team, we designed a solution with two layers:

  • A structured knowledge layer, using Model Context Protocol (MCP), to give AI models reliable, queryable access to platform systems
  • An automation layer of agent wrappers, which apply that knowledge in real scenarios without requiring a human trigger

These layers became the foundation for four specialised agents: Build, Deploy, Knowledge and Security, each targeting one of the pain points above.

Structuring data before it reaches the model

Infrastructure logs in their raw form were noisy and difficult to reason over. Rather than feeding this data directly to a large language model, we used MCP to provide structured, queryable access to the relevant systems. This allowed the model to retrieve exactly the data needed for diagnosis, in a format it could work with reliably.

The MCP layer made platform knowledge accessible. To act on that knowledge without human involvement, we built a custom MCP client for each use case. Each was configured with a clear problem scope, role context, and the inputs required to begin a diagnostic run.

Deploy Agent: automated root-cause analysis

When a pipeline failure occurs, the Deploy Agent now queries the relevant systems, runs a diagnosis through a large language model, and posts the root-cause analysis and recommended fix directly into the deployment interface. This is immediately visible to the developer, without the need to raise a support ticket. The agent runs continuously, including outside business hours, helping to remove time zone constraints.

Each diagnostic run is grounded in four specific inputs: the environment, the service, the pipeline context and the confirmed failure condition. The agent only runs when a real failure is present, and the same failure in the same context now produces the same structured result. The agent identifies the cause and suggests next steps for the developer. It doesn’t modify application code.

Knowledge Agent: developer-facing observability

We exposed the MCP configuration so that developers could connect it to the AI tools they already used, such as Microsoft Copilot or Claude, or any other MCP-compatible client.

With a simple configuration and access to the internal network, developers are now able to query deployment state and logs without needing specialist tools or direct platform access. All access is read-only, with enterprise identity controls applied to every request.

Build Agent: guided CI/CD configuration

Developers now describe their service, and the Build Agent generates a compliant build-and-deploy configuration using the organisation’s internal templates. This helps catch configuration issues early, before they lead to failures. It reduces new service onboarding from a 1–2 day coordination effort to just 30 minutes for teams familiar with the platform. This has helped increase onboarding volume to 100+ services per quarter, up from approximately 25 previously.

Security Agent: compliance visibility across the platform

The Security Agent provides a compliance dashboard with a single view across the organisation’s scanning and analysis tools. For platform and security teams, this creates a unified, real-time view of compliance across all services, something that previously required hours of manual effort involving multiple teams. The dashboard shows which applications have key controls in place (such as static analysis, code coverage thresholds, and vulnerability scanning) and highlights any gaps.

At deployment time, the agent can check releases against defined thresholds, block deployments pending security team approval, and notify teams of any outstanding vulnerabilities before their service goes to production.

Abstract digital network visual showing connected data points and secure connections across a technology platform, representing AI-powered automation, platform engineering, compliance monitoring and enterprise systems integration.

Results

Faster onboarding, stronger compliance and more engineering capacity

The impact of the solution was reflected in both faster issue resolution and a significant reduction in reliance on the specialist team.

Before these agents were introduced, much of the specialist team’s time was spent on reactive work: diagnosing pipeline failures, fielding onboarding questions, and responding to requests for compliance information.

This work has now been largely automated. Developers receive a diagnosis at the point of failure without needing to raise a ticket. New service configurations can be generated in under an hour, replacing what had previously been a multi-day coordination effort and helping to increase platform onboarding significantly. And compliance information is now visible to the teams who need it, without the need for anyone to compile it.

As a result, developers are able to resolve issues more independently, and the specialist team is able to focus on improving and evolving the platform, rather than responding to day-to-day queries.

If you’re looking to accelerate platform onboarding, strengthen compliance visibility and enable teams to work more independently, we’d be happy to explore how a similar approach could work in your organisation.

You may also like

Blog

Experimenting to enabling: How to think about AI when platform engineering

Blog

Increasing delivery throughput with self-service tooling, process, and documentation

Blog

Unlocking AI Innovation with Adaptive Guard Rails

Get in touch

Want to know more?

Are you interested in this project? Or do you have one just like it? Get in touch. We’d love to tell you more about it.