| TL;DR: Enterprises can see exactly how their systems are performing, but not whether those systems survive a failure they have never faced. Dynatrace tracks the first. Gremlin tests the second, running controlled failures to see how a service responds when a dependency dies or a cloud zone goes down. Gremlin has now launched a native app inside Dynatrace, available in the Dynatrace Hub. Engineers can now run tests without switching tools, and for the first time, reliability scores appear next to the metrics executives already track. |
Many engineering teams open a dashboard every morning. On observability platforms, that dashboard watches live systems and reports back with things like: this service is responding in 340 milliseconds, this one is throwing errors, this dependency slowed checkout at 2 am. It is useful, and enterprises have spent years building these systems around that visibility.
However, every one of those numbers reports something that has already finished happening. Kolton Andrus, CEO and co-founder of Gremlin, puts it directly: “all that data is backward-looking, meaning it reports what already happened.”
So teams read the history and extrapolate. And, usually, that works until the failure is one the system has never faced. For instance, an AWS region goes down, or a payment provider stops responding. Suddenly, the past is not much help in telling you what happens next.
Gremlin does the inverse. It causes failure on purpose, under controlled conditions, and watches what the system does. Whatever survives has been tested. Whatever does not becomes a risk the team can address before a customer discovers it.
That is what Gremlin set out to fix.
When Confidence in Resilience Is Not Enough
The cost of not knowing how systems hold up under failure became very visible in October 2025. The AWS US-EAST-1 disruption affected businesses worldwide, and Parametrix estimated losses to US companies between $500 million and $650 million.
Andrus says the outage exposed a painful gap for companies that believed they had designed sufficient resilience into their systems.

That is the uncomfortable part. Resilience is one of those infrastructure investments that can remain unverified until the exact moment it is needed. And by then, the outage is already happening.
The Tool Sprawl Problem
For organizations that wanted to answer the resilience question proactively, the workflow has often been fragmented across different tools. Reliability testing happened on one platform whereas observability data lived in another. Results had to be manually reconciled across both.
In practice, that creates more than a workflow problem. As Andrus puts it, “tool sprawl is a real problem in any enterprise company, and it is easy for signals to get lost while jumping between tools”.
Gremlin had actually built its own standalone dashboards to address part of this. They can be customized to the service, team, and company levels, with access controlled through role-based account controls. Executives at enterprise customers used them to get a high-level view of reliability. Individual teams used them to dig into scores when something dipped.
The conversation shifted at Dynatrace earlier this year, when Gremlin started talking to the Dynatrace team about their app integrations.

Rather than asking teams to change where they work, the idea was to bring resilience testing into a workflow they already use. Andrus frames it as expanding the menu rather than replacing it. “It’s less of a conversation about a Dynatrace dashboard vs. a Gremlin dashboard and more about giving customers more options.”
The App Gremlin Built Inside Dynatrace
The result is a native Gremlin app inside Dynatrace, available now in the Dynatrace Hub.
It lets engineers run resilience tests and track results against Dynatrace’s existing metrics, including response time, requests per minute, and failure rate. Gremlin’s reliability scores surface directly inside the dashboards engineers and executives already use.
Existing Dynatrace alerts and events serve as health checks and automated halt conditions. These tests stop automatically if a service moves outside defined thresholds. The app supports continuous reliability testing, disaster recovery validation, risk detection, and dependency mapping.
It works across bare metal, on-premises, multi-cloud, and serverless environments. Its production safety controls, including blast radius management and automatic rollback, apply throughout.
What Changes When Executives Can See the Score
The more significant shift may be less about the engineering workflow and more about what becomes visible to leadership for the first time.
Gremlin’s reliability scores are built around how a system responds to simulated failure conditions, such as cloud outages, dependency issues, and Kubernetes errors. When a score goes up, the system has proven resilience to more tested failure scenarios. When it falls, the platform interprets that as greater exposure to those failure conditions.
Until now, those scores lived in a separate environment that executives had to log into deliberately. Bringing them into Dynatrace changes more than the location of the data.

That is the real change on the screen. Past, present, and future now sit next to each other for the first time. An executive who wants to know whether the organization can withstand a regional cloud disruption no longer has to wait for one to happen. They can simulate the failure condition and verify the answer inside the dashboard where they track everything else.
Andrus shares that it also changes what reliability spending has to justify. Instead of tracking past performance and extrapolating from it, leadership gets a direct measure of whether the money worked.
The Banks Did Not Just Buy Gremlin, They Shaped It
Four of the five largest US banks use the platform. The more interesting part is not that they bought it, but what they pushed Gremlin to build.
In financial services, outages carry regulatory consequences in addition to direct financial ones. Configuring mechanisms such as autoscaling or multi-zone redundancy is only part of the job. As Andrus puts it, “it isn’t enough to just set up autoscaling or multi-zone redundancy and hope for the best.” These institutions have to prove those mechanisms work before an incident.
That requirement shaped the product. Some of these banks have been customers since Gremlin’s early days, and Andrus credits their methodical approach with driving the company’s emphasis on testing safely.
That meant multiple layers of fail-safes, rollback technology that returns applications to their previous state the moment anything goes wrong, role-based authorization, and enterprise reporting. In other words, a tool that deliberately breaks production systems became safe enough for a bank because banks insisted on it.
The Habit Matters More Than the Tool
Gremlin cites customer outcomes including a 50% reduction in downtime, a 90% reduction in disaster recovery testing time, and 99.99% availability on a new platform migration.
However, Andrus is unusually direct about what actually produces those numbers, and it is not simply buying the tool.

Gremlin runs its own suite of tests every week on production systems, then reviews the results and reliability scores in engineering meetings. If a score drops, the team digs in, finds the risk, fixes it, and tests again to verify the fix. All of it happens before any customer gets impacted. Andrus describes it as a light lift folded into the on-call rotation, and credits it with keeping Gremlin at five nines of availability.
He is also clear that this does not require a company-wide cultural overhaul on day one. In fact, many of Gremlin’s most successful customers started small with critical systems, proved the approach worked, then expanded outward.
The Dynatrace integration is designed to make that habit easier to build. “In many organizations, that muscle has already been built for observability,” Andrus says. Teams already review dashboards in meetings and sprint planning. So, the integration removes the separate login and interface, allowing reliability testing to become part of a routine that already exists.
Why This Move Matters Beyond the Integration
Gremlin already had dashboards that executives were using. Still, it chose to go live inside someone else’s interface anyway.
Gremlin has built integrations with Datadog, Grafana, and AWS over the years, but the Dynatrace app goes further. Instead of simply passing data between platforms, resilience testing now becomes part of the environment where engineers already monitor system performance and review incidents.
That also reflects something Andrus returned to throughout our conversation. Resilience improves when testing becomes a regular habit rather than an exercise reserved for disaster recovery planning or the aftermath of an outage.
The Dynatrace app makes that habit easier to build. Engineers can test failures, see the results alongside the metrics they already follow, fix what does not hold up, and test again without moving between different tools.
For Gremlin, that is the bigger opportunity behind the integration: making resilience testing part of how enterprises operate every day, rather than something they think about only when a system goes down.





