Hackers Weaponize AI Prompt Injections in Phishing Campaigns

Hackers Weaponize AI Prompt Injections in Phishing Campaigns

A carefully placed line of invisible text inside an email can now influence both a person and the AI assistant that screens the inbox, turning a routine phishing message into a two-layer attack. Barracuda researchers found that attackers had begun embedding prompt injections in messages meant for users, hoping to steer inbox tools that summarize mail into describing a malicious note as urgent, legitimate, or high priority. That change mattered because a summary often appeared before the original message was reviewed, giving the attacker a head start and making the message seem safer than it really was. In a phishing flow built around pressure and speed, even a small shift in presentation can decide whether a target opens an attachment, follows a link, or pauses long enough to question it.

To keep the instructions hidden from the human eye while leaving them visible to machine models, the campaigns used HTML comments, white-on-white text, Base64-encoded data, and zero-width characters. Some of the messages leaned on public-sector domains or convincing internal language so they could pass reputation-based screening without drawing attention. Others pointed to password-protected attachments that could deliver malware or harvest credentials once opened. The same technique also showed up in business contexts outside email, including fake payment updates that could alter vendor banking details and poisoned web documentation that could push coding assistants toward unsafe snippets. The pattern was clear: attackers were no longer aiming only at the recipient. They were aiming at the software that helped the recipient decide what mattered.

Why The Trick Worked

The method worked because it did not replace classic phishing; it layered on top of it. The message still borrowed familiar cues, such as vendor names, internal tone, and a sense of urgency, but the hidden prompt was aimed at the machine that processed the inbox first. Once that assistant treated the email as context, it could be pushed into downplaying warnings, elevating a false request, or ignoring the safety rules that would normally keep sensitive actions out of reach. That made a single email more dangerous than it looked. A forged summary could encourage a wire transfer review, expose data that should have stayed private, or nudge a user toward opening a file that contained malware. Reputation filters were never built for that kind of blended manipulation.

The safest response had treated AI as another untrusted input stream, not as a supervisor. Security teams that stripped comments, encoded text, and non-printable characters before content reached a summarizer reduced the chance of hidden instructions surviving the trip. They also kept financial approvals in human hands, separated external text from system prompts, and watched for repeated injection attempts across mailbox traffic and documentation systems. That approach had not eliminated phishing, but it had made the attack much harder to automate at scale. The practical lesson was simple: as AI became part of daily workflow, defenders had to sanitize content, isolate instructions from data, and verify any action with real business impact before it moved forward.

Subscribe to our weekly news digest.

Join now and become a part of our fast-growing community.

Invalid Email Address
Thanks for Subscribing!
We'll be sending you our best soon!
Something went wrong, please try again later