Microsoft has introduced a new capability in Defender for Office 365 to protect against prompt injection attacks, which target AI-powered email workflows, such as Microsoft 365 Copilot.
This update reflects the evolving threat landscape, where attackers increasingly attempt to manipulate AI systems instead of directly deceiving human users.
Prompt injection attacks involve embedding malicious instructions within email content that an AI assistant processes. Rather than using traditional phishing tactics like fake links or urgent messages, attackers craft text designed to influence how an AI model interprets and responds to a communication.
These malicious instructions can be found in the email body, subject line, attachments, or even hidden elements like invisible text or encoded content.
For instance, a harmful email might contain a concealed directive instructing an AI assistant to mark the message as safe or to forward sensitive information to an external address.
If the AI follows these instructions, it could lead to data leakage, incorrect threat classifications, or unintended automated actions.
Microsoft Defender for Office 365 Prompt Protection
Microsoft Defender for Office 365 mitigates this risk by detecting prompt injection attempts during the email filtering process, before the messages reach end users or AI systems.
This protection is automatically enabled and integrated into the existing mail flow, meaning organizations do not need additional configuration to utilize it. The detection mechanism combines large language model analysis with traditional email security signals.
Defender evaluates the complete structure of incoming messages, including visible content, HTML markup, hidden text, quoted replies, and attachments. It also normalizes obfuscated or encoded content to ensure that concealed instructions are effectively analyzed.
When a prompt injection attempt is detected, the message is categorized as “high confidence phishing,” with a specific label indicating prompt injection.
Security teams can investigate these detections using tools such as Threat Explorer and Advanced Hunting within Microsoft Defender XDR, providing deeper visibility and correlation across incidents.
This capability highlights a key distinction between prompt injection and conventional phishing. While phishing aims to deceive human behavior, prompt injection targets the decision-making logic of AI models.
The malicious payload is no longer just a harmful link or attachment but a set of instructions designed to bypass the system’s intended behavior. Microsoft positions this feature as part of a broader defense-in-depth strategy for securing AI-driven environments.
While AI applications like Copilot have built-in safeguards such as input validation, prompt isolation, and output filtering, Defender for Office 365 adds an early layer of protection at the email gateway. This ensures that malicious content is blocked before any AI system, including third-party tools or custom automation, can process it.
The introduction of prompt injection detection underscores the importance of adapting security controls to emerging AI threats. As organizations continue to integrate AI into their daily workflows, it becomes critical to protect these systems from manipulation.
By extending email security to encompass AI-targeted attacks, Microsoft aims to reduce the risk of adversaries exploiting one of the newest and most rapidly evolving attack surfaces.