The U.S. Department of Defense is exploring the adoption of Elon Musk's artificial intelligence system, Grok, as a replacement for the Anthropic-built Claude chatbot currently embedded across its operations. The proposed shift, reported by the Wall Street Journal, has triggered internal warnings about Grok's security vulnerabilities and its performance relative to competing AI models.
According to multiple officials who spoke with the WSJ on condition of anonymity, Grok is considered more susceptible to "data poisoning"—a technique where malicious or erroneous information corrupts an AI model's foundational training data. This vulnerability poses significant cybersecurity risks for a defense entity like the Pentagon, where data integrity is paramount.
Concerns about Grok extend beyond technical flaws. Federal insiders have described the system as overly sycophantic and easily manipulated, issues that have reached the desk of Ed Forst, the head of the General Services Administration (GSA), which oversees federal procurement. The GSA's reservations add a layer of bureaucratic friction to the proposed transition.
Why the Pentagon Prefers Claude
Until recently, military officials favored Claude over Grok, citing superior performance across capabilities critical to defense operations. Gregory Allen, a senior AI adviser at the Center for Strategic and International Studies, told the WSJ, "I do not believe they are peers in performance right now across all of the capabilities that matter to a customer like the Department of [Defense]." This assessment underscores the gap between the two systems in areas such as reasoning, accuracy, and reliability.
The preference for Claude, however, hit a roadblock when Anthropic refused to remove two key ethical guardrails requested by the Pentagon. This refusal prompted the search for an alternative, leading to Grok despite its known shortcomings.
Complicating matters further, Sam Altman, CEO of OpenAI—Anthropic's rival—signaled this week that his company would uphold a similar ethical red line, refusing to compromise on safety measures for defense use. This stance limits the Pentagon's options, as Google and Microsoft have not yet indicated a willingness to cross the line that Anthropic and OpenAI have drawn.
Grok's Track Record and Reputation
Grok, developed by Musk's xAI, has already been deployed in select parts of the federal government, including some defense applications. However, its reputation has been marred by erratic and sometimes offensive outputs, and it scores notably lower on AI benchmark tests compared to leading competitors. These issues have fueled skepticism among federal insiders about its suitability for high-stakes environments.
Musk's familiarity with federal operations, having spent much of 2025 advising on government efficiency, has not alleviated these concerns. The WSJ's reporting suggests that the push to adopt Grok may be driven more by political alignment than by technical merit, raising questions about the decision-making process.
As the Pentagon weighs its options, the standoff between ethical safeguards and operational needs remains unresolved. Without a viable alternative that meets both security and performance criteria, the department may be forced to proceed with Grok, accepting the associated risks.