SoS Logo

An Interview with Bartley Richardson About Building CrowdStrike SafeMind

An in-depth discussion about CrowdStrike's data advantage, SafeMind's architecture, model training, harness engineering, and digital twins.

Oct 8
Public Companies
an-interview-with-bartley-richardson-about-building-crowdstrike-safemind

I recently had the opportunity to talk with Bartley Richardson, the Chief AI and Autonomous Systems Officer at CrowdStrike.

Bartley leads CrowdStrike's cybersecurity AI initiatives. He's the driving force behind their Cyber Superintelligence Lab, SafeMind, and more.

This discussion was a chance to go beyond the Fal.Con keynotes and dive deeper into several AI-related topics.

If you want to understand where the frontier of cybersecurity-focused AI is going, Bartley Richardson is the first person you should listen to.

Note: This interview has been lightly edited for clarity.


CrowdStrike’s data advantage and model training

Cole Grolmus: CrowdStrike has said many times that it has a significant data advantage for training cybersecurity-focused models due to the amount of data that’s been collected and annotated over time.

How much of an advantage does that give you over anybody else in the world who either has less cybersecurity-related data or none at all? Is there a point of diminishing returns? Or is it infinite in terms of how it improves post-training models?

Bartley Richardson: I hesitate to say anything is ever infinite. There would be a point where you say, “that is enough.”

When we’re using reinforcement learning on the Blue Solano side, we want to make sure it captures the ground truth of how a defender would approach a problem.

We don’t want to say, “here’s the input, and there’s the output.” We want to capture defenders’ reasoning processes and reasoning traces and put that in there.

You have to be careful about the data blend you use for training. There are many different scenarios, everything from very simple to moderate to complex to nation-state.

We crank up the data blend on how an advanced adversary, human or AI, does reconnaissance and spreads out and fans out into your network.

The reason we crank up that part of the data blend is that, in the reinforcement-learning process, you can generalize based on it. I don’t want a model that can only produce exactly what it saw. I want it to generalize based on everything it has seen. The more complex scenarios it sees, the more it can reason.

It’s not just about feeding it all the data. The data blend is the secret sauce, and that’s part science and part art. We run multiple experiments with reinforcement learning, then adjust the data blend.

Cole Grolmus: That’s interesting because it’s a different answer than I expected, in a good way.

People have heard and talked a lot about CrowdStrike’s data advantage, which is both real and significant.

Part of what you just talked about is net-new data learned in the lab through continuously running these models against each other and experimenting.

That’s separate in some ways from the data advantage that has been discussed.

Bartley Richardson: It’s separate in some ways, but it also fits into CrowdStrike being a net data creator from the beginning. We’re creating new data. The modality is a little different than it was before.

The prior modality was that I have a bunch of threat researchers, defenders, and other people logging Jira tickets, and it’s annotated. Now we’re doing something similar in this agentic world and capturing it. It’s still net data creation in the platform.

The cool thing is that something like SafeMind is purpose-built around detecting and remediating in the running environment.

But let’s take another example, like identity. We didn’t share much about it, but we’re working on other similar things with red and blue. In that case, I select a completely different data blend.

Depending on what you’re trying to do, I look at all these dials. I might crank this dial up a little and turn that dial down a little. Then you perform successive iterations so you get an accurate result for a different set of use cases like identity.

When we say CrowdStrike has the best data, the nuance is that we have the best data because we have so much heterogeneity in it. The heterogeneity of the labeling lets us go down multiple paths.

SafeMind’s architecture and harness engineering

Cole Grolmus: Is focus on real use cases and workflows the difference that will make everything you built with SafeMind — the models, the harness, all of it — successful in the field?

As opposed to the last four years of cyber models that were model-only and never gained a meaningful amount of traction.

It feels like there is a level of practicality involved that will take SafeMind from a lab experiment to something a practitioner can use.

Bartley Richardson: The reality today is that there are multiple models in the Red Tempest family and multiple models in Blue Solano, plus multiple harnesses around Red Tempest and multiple harnesses around Blue Solano. Then there’s the multi-agent architecture.

What makes that successful is that we have customized models, which are actually a family of models, and customized harnesses. They could be multi-agent or multi-swarm harnesses.

Some could have different purposes, such as reconnaissance. We might have another harness that’s the best in the world at working with specific platforms. It understands how to operate that platform like an operator would. Another harness might understand how to do sophisticated, deep technical research on security vulnerabilities across GitHub, X, and the web.

These purpose-built harnesses are built along with the model. That coordination makes them work. You can take a model and a harness, put them together, and be okay.

But what you really want is to put the model and harness together, run them, and use the harness traces. You train to align the model so it works better in the harness. That happens for us at the post-training and reinforcement-learning stages. We also modify the harness as the model evolves, ensuring the model is able to effectively drive it towards the outcomes we want.

So, part of what makes SafeMind effective is the models, but it’s also: “I want to design this model with this harness.”

Cole Grolmus: Model training obviously matters, and I don’t want to diminish that part. But it seems like your harness engineering has been underestimated so far.

When we talk about SafeMind’s harness, are we truly talking about a harness, or are there other agentic factors like skills, core agent files, etc.?

I know technically that’s all the harness, but I’m trying to break down the subcomponents that make this work.

Bartley Richardson: You’re right. The harness incorporates the ability to work with tools and the ability to import and execute skills.

When you have a harness connected to tools or importing skills, the model gives the harness a directive. You might give it a task such as, “Fetch the last six hours of Falcon telemetry from this host.” Then it’s up to the harness, through a skill or another deterministic mechanism, to execute that task.

How well the harness does the task depends on how well the model inside it understands the harness, tools, API calls, and everything else it needs to make a valid call. We see that manifest in interesting ways.

There’s a fuzzy line. You don’t have to fully optimize it. You can let the agent try multiple times. For example, take a structured query language (SQL). With SQL, you know whether it got it right. It works or it doesn’t. There are costs associated with that, or you can try to one-shot it.

There are pros and cons to each. If you care about speed, latency, and cost, you want to one-shot it. If it’s a longer task and you don’t care as much, you can spend less time training the model, which also costs money, and let the agent iterate two or three times to get the skill calls right.

It’s like when you have a newborn. They have a brain and a body. Humans are one of the only animals born unable to do anything. It’s amazing we survive.

The newborn has a brain and it’s evolving, but they don’t understand how to use the brain with the body. There’s no walking motion. They take steps and do other things. As they figure out how to do this — what does balance feel like, what can I grab onto — that goes back into the brain, which learns and rewires. The brain becomes better at coordination and balance. It gets better at recognizing, “That’s hot; don’t touch it.”

This is a toy example, but it’s essentially what we do with harnesses and models: make the brain better at working with the harness.

Cole Grolmus: Building on the practicality point, I want to scope what the defensive model can do.

I think it’s clear what the red-team model does with offensive security. What’s less clear from a practitioner’s point of view is the scope of what the defensive model will be able to do.

As someone who spent a decade in big enterprises watching how things get done, a lot of the workflows are part process and part interfaces with other apps and infrastructure.

How far will that go? How far does it go today, and how far do you want it to go?

Bartley Richardson: Right now, you can get it in two different flavors.

One is findings. In the easiest case, we can prepare a prioritized list of actions you need to take. That’s still helpful.

Even if they don’t have the Falcon platform fully deployed, we want to help. We can provide a list: “Here are the remediations you might want to look at. Go audit your IAM rules, rotate these credentials, and do these other things.”

We can do much more when customers have the Falcon platform fully deployed. If you have more of the Falcon platform, such as Falcon for IT, we can autonomously close the loop with this data.

We envision a world where auto-remediations happen in the platform — at every endpoint, sensor, and instrumentation point. Instead of giving you a list of what needs to be done, the actions get sent to your MCP server or wherever else they need to go, with auditability around them. Touching these other systems is on the roadmap.

Cole Grolmus: Did you follow a conceptually similar approach for harness engineering as you did with the models where NVIDIA Nemotron was the base?

Did you start with something like Pi and build on top of it, or did you build the harness from the ground up?

Bartley Richardson: For counter-adversary operations and certain parts of the remediation chain, we built those harnesses ourselves from scratch.

Jensen Huang said something to the effect that people should use off-the-shelf models wherever possible. The same is true with harnesses. I’m going to choose something like Pi for most of my cases — it doesn’t have to be Pi, any open source harness.

For the more bespoke cases, it turns out the codebase can be incredibly small because I don’t need it to do everything. I need it to be a singular expert at a specific task, such as malware deconstruction. It’s not too difficult to write, but you do have to spend time verifying the harness.

Cole Grolmus: I could see a world where the SafeMind models and harness could be offered as a standalone product.

It’s really good at finding and fixing security issues, whether you do that through CrowdStrike or not.

If you want to connect it directly to Salesforce and fix a bad Salesforce configuration, great. It’s good at executing security-related jobs no matter where the work needs to get done.

Bartley Richardson: That’s how we’re thinking about it. SafeMind is a complete platform. One way to consume it will be in the Falcon platform, but it doesn’t have to be consumed that way. You could consume it from a different platform.

When we launched Project QuiltWorks, the vision was, “Let’s establish a set of trusted partners.” We know we’re going to make highly capable models. Let’s make sure trusted partners have a vehicle to get access. For me, that’s the key part of the Project QuiltWorks program.

I think about it this way: you’re in the Project QuiltWorks program. I trust you to have this model. I trust you to keep it safe and secure, and I want you to have the capability to put it in your own harness for a use case bespoke to you.

Certain companies might have bespoke cybersecurity use cases and workflows. We want to enable that and trust you to keep it safe.

Reciprocally, how do we make these models better? We look at the traces, especially when there’s unexpected behavior. We’ll trust you to have and secure the model, and we’d love you to trust us by sharing those traces back in a way that preserves privacy.

How digital twins work

Cole Grolmus: Could you explain more about the need for digital twins and how they work?

I understand the concept, but as someone who spent a long time in complex enterprise IT environments, I don’t know how you replicate an environment like that with reasonable accuracy. You have SAP, Active Directory, mainframes, and everything else.

I would love to hear more about how this works.

Bartley Richardson: It’s pretty slick. You can instantiate it in several ways.

The easiest relies on a series of models. You can prompt it in natural language: “I’m a 500-person enterprise company. I use these applications, have these types of endpoints, and see these types of data.”

Over several minutes, depending on your compute, actual VMs get loaded. You can have a repository of operating-system versions, application-system versions, sample behavioral network configurations, and all of that. The VMs get spun up, populated, and networked together.

You can also feed it network or architecture diagrams. At CrowdStrike, we do all those things. We also use it when our red teamers do offensive work and penetration tests. We use the same environment to validate those tests.

They document everything they do in Jira tickets. We take the raw Jira tickets and put them into the LLM that drives the front end. It spins up the environments, runs the attack, matches their pattern, and does all those things.

Another way to do it is if you’re connected to the CrowdStrike Falcon platform.

We pull a ton of data from Falcon for IT, Falcon Exposure Management, and other sources. We know which versions you’re running and what CVEs might be present. We can recreate an environment with, for example, 10 Windows hosts, including two that have CVE-XXX.

There’s no magic behind it. We’ve created an interesting set of models and a fairly sophisticated LLM prompt-based interface. You can chat with it or feed it unstructured data, and out pop these digital environments.

We want successive iterations in a digital-twin environment. Then we can do things like create a digital twin of a newly remediated environment and ask whether parts that weren’t perfect in the twin now matter.

What we don’t want to do is take so much time crafting the perfect replica that we lose time, because it can’t be perfect.

This is back-of-the-envelope math, but it bears out in our testing: you can get upward of 80 to 90 percent with this digital product.

Many digital-twin efforts have fallen short of the promise. They try to be perfect. They let perfect get in the way of good.

Share Article

Related Articles

SOS Logo

Cybersecurity, clarified

Strategic intelligence on the cybersecurity ecosystem, straight to your inbox.