Centre for Police Technology Launches a 31-Day Cybersecurity Knowledge Series for Police, LEAs, Corporate Investigators, Digital Forensics, Fraud, Cyber Risk and Security Professionals
October 3, 2026: Artificial intelligence is becoming central to cybersecurity, investigations and business operations. But the same systems can also be manipulated, deceived or exploited by attackers.
This creates a growing security challenge:
Adversarial AI — attacks designed to manipulate AI systems, exploit their weaknesses or influence their outputs.
Following Day 1’s focus on Agentic AI Security and Day 2’s focus on AI Threat Detection, Day 3 of the Centre for Police Technology (CPT) 31-Day Cybersecurity Knowledge Series examines how attackers target AI and how organisations can defend against these threats.
What Is Adversarial AI?
Adversarial AI refers to techniques used to manipulate or exploit artificial intelligence systems.
These attacks can target AI models, their training data, inputs or connected tools.
Common examples include:
- Adversarial examples: Small changes to an input that cause an AI model to make an incorrect prediction.
- Data poisoning: Manipulating training data to influence model behaviour.
- Prompt injection: Malicious instructions designed to manipulate an AI model or agent.
- Model evasion: Changing malicious content so AI-based security systems fail to recognise it.
- Model extraction: Attempting to reproduce a model through repeated queries.
The danger is that an AI system can appear to work normally while producing manipulated or unreliable results.
How Are Attackers Exploiting AI?
Attackers can target AI throughout its lifecycle.
For example, criminals could attempt to manipulate an AI-powered fraud detection system so suspicious activity appears legitimate. In generative AI applications, prompt injection can attempt to make an AI assistant reveal information, misuse connected tools or ignore its intended instructions.
Another emerging threat is indirect prompt injection, where malicious instructions are hidden inside websites, documents, emails or other content that an AI system later processes.
This creates a fundamental problem:
The AI may not know whether the information it is reading is data — or an instruction from an attacker.
Why Is Adversarial AI Difficult to Detect?
Traditional security systems often rely on known attack patterns. Adversarial attacks can instead exploit weaknesses in how AI interprets information.
An image may look normal to a person but cause an AI model to make an incorrect classification. A document may appear harmless while containing instructions designed to manipulate an AI assistant.
Organisations therefore need to ask not only:
“Is the AI working?”
but also:
“Can the AI be manipulated, and what happens if it is?”
How Can Police and LEAs Address Adversarial AI?
Police and Law Enforcement Agencies may encounter AI-generated content, manipulated digital evidence, synthetic identities and AI-assisted fraud during investigations.
AI can help analyse large volumes of digital information, but investigators should:
- Verify AI-generated findings against independent evidence.
- Preserve original files, metadata and relevant logs.
- Document the tools and methods used.
- Check whether data or model outputs may have been manipulated.
- Follow established forensic and evidence-handling procedures.
An AI-generated finding should be treated as an investigative lead, not automatically as proof.
Human verification remains essential.
FCRF Launches CP-FRM to Build India’s Next Generation of Fraud Risk Professionals
How Can Organisations Defend Against Adversarial AI?
Organisations deploying AI should treat models, datasets and connected applications as part of their cybersecurity environment.
Important safeguards include:
- Testing models against adversarial inputs.
- Monitoring and validating training data.
- Restricting access to models and sensitive datasets.
- Deploying prompt-injection and jailbreak protections.
- Limiting what AI agents can access or execute.
- Monitoring unusual model behaviour.
- Maintaining human approval for high-impact actions.
- Conducting continuous AI security testing and red-teaming.
The objective is not simply to prevent every attack, but also to limit the damage if an attack succeeds.
Emerging Technologies Fighting Adversarial AI
Technology companies are increasingly developing security controls specifically for AI systems and agents.
NVIDIA Open Agent Safety Platform
In September 2026, NVIDIA introduced its Open Agent Safety Platform, including OpenShell and Sentry. OpenShell provides a controlled runtime environment for AI agents, while Sentry is designed to monitor agent behaviour and enforce security policies.
Microsoft AI Gateway
Microsoft has introduced Prompt Injection Protection through its AI Gateway to help detect and block malicious prompts and jailbreak attempts targeting AI applications.
Google AI Security
Google is developing layered defences against indirect prompt injection and uses security testing and red-teaming to identify new attack techniques targeting AI applications.
OpenAI Agent Security
OpenAI is researching layered defences against prompt injection, including model training, monitoring, sandboxing, red-teaming and controls around sensitive actions.
The emerging approach can be summarised as:
Detect → Isolate → Monitor → Restrict → Verify
How Can Individuals Stay Safer?
Individuals using AI tools should avoid unnecessarily sharing passwords, financial information or confidential documents.
Users should also be cautious about giving AI agents unrestricted access to email accounts, files, financial systems or other sensitive resources.
For important actions:
Verify → Review → Approve
before allowing an AI system to send information, make a transaction or perform another consequential action.
From AI Adoption to AI Security
As AI becomes part of cybersecurity, policing, financial services and everyday applications, securing these systems is becoming as important as securing conventional IT infrastructure.
For organisations, this means testing AI before deployment and continuously monitoring it.
For investigators, it means understanding that AI-generated outputs can potentially be manipulated.
For individuals, it means treating AI as a powerful tool — not an unquestionable source of truth.
The central principle is simple:
AI must be secured against adversaries, not just used against them.
Day 3 — Adversarial AI
31 Days | 31 Key Topics | October 2026
A Cybersecurity Awareness Month Knowledge Initiative
Created by Centre for Police Technology (CPT)
Follow Centre for Police Technology (CPT) for the complete 31-Day Cybersecurity Knowledge Series.
Follow for daily updates on cybercrime, corporate fraud, DFIR, hacking, investigations, and digital forensics