The problem
The Deep Space Network is NASA’s array of large radio antennas in California, Spain, and Australia, and the link to every spacecraft operating beyond the Moon.1 A missed communication pass can mean lost science data, so the network’s most important property is availability.
That makes routine security work delicate. The standard fix for a vulnerability is to patch and reboot, and rebooting a mission-critical machine during an active tracking pass is not an option. Scanners rank findings by CVSS,3 a severity score that is deliberately context-free: it can’t know what a machine does, or whether it can go down tonight. Analysts fill that gap by hand, cross-checking every alert against documentation, configuration records, and years of institutional knowledge.
Meanwhile the number of disclosed vulnerabilities keeps climbing, and LLMs are beginning to speed up how attackers find and exploit them. Defense that depends entirely on manual correlation can’t keep pace, and correlation is exactly what a well-equipped LLM agent is good at.
What I built
An end-to-end pipeline that turns a raw scanner finding into a reviewed, structured remediation recommendation. The agent decides which tools to call to gather context, reasons about the vulnerability, and returns a validated assessment. It recommends; it never acts. Step through it below, and switch between the two example alerts.
Example alert · synthetic data
Alert
A vulnerability scanner reports a finding with a CVSS score, a universal severity number. It knows nothing about what the machine does for the Deep Space Network, or what applying the fix would interrupt.
Context
Instead of stuffing every document into the prompt, the agent calls MCP tools and fetches only what it needs: engineering documentation first, then the wiki, the acronym database, and the configuration database.
search_docs("status-web-02 role")Public status page. Redundant pair behind a load balancer.config_item("status-web-02")Non-mission system, standard maintenance window.wiki_search("package update rollback")Rolling update, one node at a time. Nothing used for tracking restarts.
search_docs("sp-host-14 role")Processes downlink signal during active tracking passes.config_item("sp-host-14")Mission-critical. No hot spare.wiki_search("kernel patch procedure")Requires a reboot. Schedule outside tracking passes.
Agent
A Pydantic AI agent has to fill a typed schema, not write prose. When its output breaks the schema, the validation error goes back to the model and it tries again, at temperature zero.
- Attempt 1Valid on the first try.
- Attempt 1
difficulty: "medium"is not one of"easy"or"hard". Error sent back to the model. - Attempt 2Valid. Passed to output.
Output
The result is validated JSON: a difficulty class based on operational feasibility, the reasoning behind it, deployment considerations, and the documents it relied on. Whether a flaw is known to be exploited comes from CISA's KEV catalog2 as a hard fact, not a model guess.
{
"vuln_id": "SYN-A",
"severity": "High",
"difficulty": "easy",
"reason": "Routine package update on a
redundant, non-mission host. Rolls back
cleanly; fits a normal maintenance window.",
"considerations": ["Update one node at a time"],
"known_exploited": false
}
{
"vuln_id": "SYN-B",
"severity": "Medium",
"difficulty": "hard",
"reason": "Needs a reboot of a mission-critical
host with no hot spare. Schedule outside
tracking passes.",
"considerations": ["Coordinate with the tracking
schedule", "Verify signal processing after reboot"],
"known_exploited": false
}
Review
Nothing is applied automatically. An analyst reviews each recommendation in a dashboard, grouped by difficulty, and decides. The system recommends; people act.
| Alert | Scanner severity | Agent difficulty |
|---|---|---|
| SYN-A | High | easy |
| SYN-B | Medium | hard |
Sorted by severity, A comes first. Sorted by what the fix costs the network, B is the one that needs planning.
The central result is that inversion. A high-severity flaw whose fix is a routine, easily rolled-back update is classified easy. A moderate one whose fix needs a reboot of core services, a firmware change, or a performance-costing patch is classified hard. The agent gets there by consulting the same internal documentation an analyst would, and it has to cite what it used.
Structured output as a contract
Free-form prose can’t feed a pipeline. The agent is built on Pydantic AI,4 which constrains the model to fill a strictly typed schema. When the output violates it, the validation error is fed back and the model corrects itself, up to a fixed number of retries. That loop is what turned a promising demo into a component other tools can rely on. Inference runs at temperature zero, and every assessment is stamped with the prompt and schema versions, the model ID, and the backend it ran on, so any result can be traced and reproduced.
| Field | What it holds |
|---|---|
severity | Standard severity level, from the scanner |
difficulty | Operational feasibility: easy or hard |
reason | Written rationale for the classification |
impact_analysis | Affected systems and mission impact |
considerations | Deployment and remediation considerations |
remediation_action_command | The exact fix, or null if it’s manual |
verification | Generated tests to confirm the patch and service health |
known_exploited, kev_* | CISA KEV enrichment, injected as fact |
Context on demand, through MCP
Putting every potentially relevant document into the prompt would be expensive and would overflow the context window. Instead, each context source sits behind a Model Context Protocol5 server, and the agent retrieves only what a given vulnerability needs. The sources are the engineering documentation (the primary, authoritative source), an internal wiki of procedures and known issues, an acronym database, a configuration-item database, and the existing vulnerability data. Each server runs as a local process that talks only to internal resources, so sensitive data stays inside approved boundaries.
The backend is model-agnostic. Switching between three approved backends takes one configuration value, so analysis can be routed by data sensitivity, and the tool improves as the underlying models do.
1 · Context sources
2 · Context connectors
3 · Reasoning
4 · Outputs
5 · Human in the loop
Caching, checkpoints, adaptive pacing, and cost tracking wrap every model call.
Chains, not lists
Attackers rarely use one vulnerability. They chain them,6 turning a minor foothold into a position from which a serious compromise becomes possible, and a flat list of findings hides that. A second agent reasons about how vulnerabilities across interdependent systems could combine, and the dashboard draws the result as an interactive attack-path diagram.
- Low severityWeb footholdA minor flaw on an exposed service gets an attacker in.
- Medium severityPrivilege escalationA second, unrelated flaw raises their access on that host.
- Enabled by the chainLateral movementDependencies between systems carry them further in.
- CriticalCritical compromiseEven though none of the individual findings looked urgent.
Built to run for real
- Caching
- Repeated context queries are cached, cutting cost on repeated queries by an estimated 50–70%.
- Checkpoints
- Progress is saved after every assessment, so an interrupted run resumes and loses at most one item.
- Adaptive pacing
- When a backend rate-limits, the delay between requests widens and honors cooldown hints, then relaxes as pressure clears.
- Cost tracking
- Input, output, and cached tokens are recorded per assessment, with actual and cloud-equivalent cost.
What comes next
The next step is measuring how good the recommendations really are, by comparing the agent’s calls with what experienced DSN analysts would decide. Doing that well takes a purpose-built dataset: real vulnerabilities from the network, each paired with an analyst’s judgment and the reasoning behind it. Assembling that takes real analyst time and care, which is why it comes after the prototype rather than alongside it. Until it exists, I’m not putting an accuracy number on the system.
Two capabilities are already prototyped and ready for the same hardening: vision, so the agent can read the diagrams in DSN documentation directly, and a tightly sandboxed container that tries to reproduce an exploit to see whether it’s actually feasible.
Two research questions came out of the internship: whether structured vulnerability-test data improves the agent’s grasp of real-world impact, and whether it can reason over firewall rule sets to prioritize by actual reachability7 instead of assuming worst-case connectivity.
References
- NASA Jet Propulsion Laboratory. Deep Space Network.
- Cybersecurity and Infrastructure Security Agency. Known Exploited Vulnerabilities Catalog.
- Forum of Incident Response and Security Teams. Common Vulnerability Scoring System (CVSS).
- Pydantic. Pydantic AI documentation.
- Anthropic. Model Context Protocol.
- Hutchins, E. M., Cloppert, M. J., & Amin, R. M. (2011). Intelligence-Driven Computer Network Defense Informed by Analysis of Adversary Campaigns and Intrusion Kill Chains. Lockheed Martin.
- The MITRE Corporation. MITRE ATT&CK.