Garbage In, Verdict Out: AI-Driven Malware Triage and the Disassembler It Never Questions
Sample volume outpaces analysts, and modern Go, Rust, Swift, and .NET binaries bury malware logic under thousands of runtime functions, most of them noise. To make first-pass triage scale, we built a static-only pipeline that pairs a headless disassembler with a frontier large language model.
It began with interactive experiments: first an analyst-in-the-loop IDA Pro LLM chat plugin for conversing with the model, then an MCP server that ran IDA Pro as a backend and let the agent drive it directly, pulling functions and cross-references on demand. It then grew into an automated pipeline that triages packed and unpacked samples across Windows, Linux, and macOS in a range of languages. The disassembler performs local extraction, filtering retains only the functions that matter, and a single model call returns a structured verdict covering the malware family, IOCs, suspicious APIs, and behaviours, all at a bounded cost.
This talk walks that system end to end: how it’s built, what it gets right, and how we keep the model honest instead of taking its confident answers on faith.
Then comes the question nobody asks: the model interrogates the malware while blindly trusting the tool that produced its evidence. If we swap that tool, does the verdict change? We put it to the test and report what we found.
Key takeaways
- How to build a static, cost-bound LLM triage pipeline that scales beyond analyst capacity and evolve it from an interactive tool into an automated system.
- How to keep the AI honest in practice and identify the failure mode that survives the safeguards.
- How to audit the disassembler underlying your AI and test whether it changes the verdict.

Jagadeesh Chandraiah – Sophos
Jagadeesh Chandraiah is a Threat Researcher at Sophos X-Ops with over 15 years of experience in malware analysis, reverse engineering, and threat intelligence. His research spans Windows, macOS, and mobile threats, with a current focus on applying AI and large language models to malware analysis and automated triage. A regular contributor to the Sophos X-Ops blog, he has presented his research at international security conferences including Virus Bulletin, AVAR, CARO Workshop, and DeepSec. He holds a Master’s degree in Computer Systems Security from the University of South Wales.