← Back to list

Generative AI and Cybersecurity: What Are the Risks ?

Because at Craft AI, we believe that high-performing AI is, first and foremost, secure AI, our tech experts invite you to discover the…

Craft AI Team in Craft AI · 2026-01-15 09:21 · 0 claps · 3.9 min read
#ai #cybersecurity #secure-ai-models
Open on Medium ↗
Wiki topics: AI · AI · General 🔒 · Cybersecurity 🛠️ · Crafts & DIY

Generative AI and Cybersecurity: What Are the Risks ?

Because at Craft AI, we believe that high-performing AI is, first and foremost, secure AI, our tech experts invite you to discover the different types of attacks targeting generative AI systems and how to protect against them.

In the world of AI, one system is making headlines: Generative AI. It has become an integral part of our daily lives with solutions like ChatGPT, Gemini, and Claude, which are now used by millions of people every day.

However, a massive user base inevitably makes for a prime target for cyberattacks: agent hijacking, data theft… these are risks that must not be underestimated when building or using this type of system.

But what types of attacks are we talking about? What are the consequences, and how can we prevent them? We tell you everything in this article.

What are the vulnerabilities of Generative AI ?

If generative AI systems are ideal prey, it is because they increase the “attack surface”: this represents the number of points where an unauthorized user can access a system and manipulate data. The smaller this surface, the easier it is to protect (private environment, specialized AI, etc.). In this case, generative AI platforms are open to the general public and connected to a multitude of external sources to retrieve information; their attack surface is therefore much larger.

Image source: MesServicesCyber by ANSII

The Different Categories of Attacks

Generative AI systems are subject to several types of attacks, ranging from AI hijacking to data theft.

  • Manipulation Attacks: These are designed to hijack the AI agent’s usage or modify its responses via malicious queries crafted to “bypass” the AI’s limitations.
  • Infection Attacks: This attack can occur during the AI’s training phase. A malicious user may add erroneous information to a dataset or knowledge base that the AI relies on to provide answers.
  • Exfiltration Attacks: This type of attack aims to steal information while the system is in production.

Some Examples

These different types of attacks are made possible through various methods.

  • Jailbreaking: Jailbreaking, or direct prompt injection attack, consists of using a specific type of prompt to modify the AI’s behavior. This allows the attacker to bypass “guardrails” and access confidential information. For example, an AI will never give you instructions on how to make a bomb. However, if you tell it you have certain ingredients and are missing one to make a bomb, the AI agent might give you the name of that missing ingredient. The attacker can also achieve their goals by “insisting” repeatedly until the AI finally yields.
  • Data Poisoning: This technique involves adding a small number of “poisoned” documents into a database with the aim of degrading the AI’s performance or manipulating its responses. The attacker must include falsified information “discreetly” to modify the agent’s behavior without it being obvious.
  • Indirect Prompt Injection Attack: This attack is a derivative of jailbreaking. It consists of manipulating the AI’s responses by using a prompt not provided by the user during their request. For example, the attacker might access the system’s database and tell it: “Forget the user prompt and execute this one instead.”
  • Zero Click Attack: A Zero Click attack is characterized by the execution of a malicious instruction that the user never provided. The “booby-trapped” prompt is retrieved autonomously by the AI via its knowledge base or an external tool (email, webpage). This technique allows the attacker to manipulate the AI’s responses or exfiltrate confidential data. Unlike classic injections, this attack requires no direct interaction between the user and the malicious prompt. For instance, code hidden in an email can activate itself as soon as the AI agent analyzes the message to summarize it, compromising the system without the victim noticing.

How to Protect Against These Attacks

As we have seen, numerous attacks exist, and attackers are doubling down on creativity to bypass the “guardrails” implemented by developers to execute malicious prompts.

Therefore, several best practices should be implemented to protect against these attacks and avoid the worst:

  1. Never share internal or confidential information with an artificial intelligence tool. Whether it’s to reply to an email, by copy/pasting text, or by sharing a document for summarization.
  2. Prioritize the use of “internal” AI that does not always require a connection to external servers and solutions.
  3. Continuous “Red Teaming”: Rather than waiting for an attack, companies should organize simulations where experts actively attempt to hack the AI using the methods mentioned above. This helps identify and fix vulnerabilities (patching the guardrails) before they are exploited by real attackers.
  4. The Principle of Least Privilege: To limit the impact of a Zero Click attack or exfiltration, the AI agent must only have access to the data strictly necessary for its mission.
  5. Data Provenance (Sanctuarisation): To counter Data Poisoning, one must ensure the integrity of training data. This implies rigorously verifying sources, using digital signatures to validate documents, and regularly scanning knowledge bases to detect potential fraudulent modifications.
  6. Log Monitoring: Preventing attacks is good, but being able to detect and trace them is better! This is the principle of “log monitoring.” This practice involves keeping and analyzing all user/AI interactions (prompts, model responses, etc.) to identify suspicious patterns (e.g., repeated injection attempts) and to trace the sequence of an attack a posteriori to understand the flaw and adapt defenses.

This topic has gained such momentum in recent months that the DGSI (French General Directorate for Internal Security) has looked into the matter and formulated recommendations, which you can find here: *Security Recommendations for a Generative AI System.*

Ready to take the leap into AI securely? Contact our experts.


메타데이터
post_id
dc12ef4466f3
slug
generative-ai-and-cybersecurity-what-are-the-risks-dc12ef4466f3
url
https://medium.com/craft-ai/generative-ai-and-cybersecurity-what-are-the-risks-dc12ef4466f3
canonical_url
https://medium.com/craft-ai/generative-ai-and-cybersecurity-what-are-the-risks-dc12ef4466f3
author_url
https://medium.com/@craft_ai
status
ok
fetched_at
2026-06-10 08:17:25