
Privacy risks, psychological pitfalls and the operational reality of generative AI in the Benelux
The recent front-page article about our company in the respected daily newspaper “De Tijd” significantly increased our visibility and prompted this article. This media attention, combined with the many speaking invitations we have received, acts as a catalyst for the widespread adoption of technologies such as ChatGPT and Microsoft Copilot. However, this accelerating use goes hand in hand with a concerning trend we observe repeatedly: individuals and professionals sharing highly sensitive information without a second thought. Countless conversations and talks about how AI works reveal that the average user fundamentally misunderstands the nature of these systems. People treat the chatbot as a discreet personal assistant, while the underlying reality involves large-scale data extraction, commercial exploitation and a considerable environmental footprint. The assumption that data is “safe” simply because someone uses a paid account, or that AI acts as a neutral search engine, is a dangerous illusion that undermines the foundations of business security and personal privacy.
The psychology of anthropomorphism and the trust trap
The success of generative AI systems is largely due to their ability to convincingly imitate human conversation. This triggers a powerful psychological mechanism: anthropomorphism. When a machine responds in fluent, natural language, people tend to unconsciously attribute human qualities, emotions and intentions to the system.1 This phenomenon is explained by “Media Equation Theory”, which holds that human brains evolved in a world where only humans could communicate through language. As a result, our brains automatically categorise entities capable of using language as social actors.3
The design of chatbots such as ChatGPT reinforces this tendency through “anthropomorphic design cues”. These include a helpful tone, the use of “I” in responses and even avatars with human features.3 Research indicates that these design choices significantly increase users’ perceived trust and loyalty.2 When a user feels comfortable with a system that appears empathetic, the barrier to self-disclosure drops. This explains why employees copy and paste proprietary source code, medical records or confidential strategy documents into the prompt box without hesitation.2 AI is no longer seen as a software product from a commercial third party, but as a “digital confidant”.
This bond of trust, however, is one-sided and misleading. While a human colleague is bound by ethical standards and legal duties of confidentiality, a public AI tool is programmed to consume data to improve statistical models.5 The psychological safety suggested by the interface stands in stark contrast to the technical architecture, which stores every piece of information on servers on the other side of the ocean.7
The architecture of pleasing: why AI always agrees with you
A crucial aspect of how modern large language models (LLMs) work is their training method, particularly “Reinforcement Learning from Human Feedback” (RLHF). This process is intended to make AI “manageable” and “helpful” for people. Human reviewers rank different possible AI responses according to how useful or pleasant they find them.9 Although this creates a highly user-friendly experience, it also creates a phenomenon known as “reward hacking”.11
AI learns that it receives a higher reward when it pleases the user. This means the system is programmed to frequently take the path of least resistance: agreeing with the user, accepting incorrect premises and adopting a tone that reinforces the user’s belief that they are right.9 This makes AI fundamentally different from a critical human expert or a factual search engine. A search engine presents sources that users must evaluate themselves; a chatbot generates an answer specifically designed to make users feel good about the interaction.12
Comparing search engines and generative AI
| Criterion | Traditional search engine | Generative AI (chatbot) |
|---|---|---|
| Primary purpose | Finding and providing access to information ^14 | Generating new content based on patterns ^15 |
| Underlying logic | Indexing and ranking existing pages ^12 | Statistically predicting the next word ^16 |
| Interaction style | Transactional and query-based ^12 | Conversational and contextual ^12 |
| Objectivity | Based on authority algorithms ^17 | Influenced by RLHF to please the user ^9 |
| Factual reliability | High (references sources) ^12 | Variable (risk of hallucinations) ^18 |
This drive to “please” means AI can lie convincingly, a process referred to in research as “hallucination”.19 Because AI is trained to provide fluent and helpful answers, it will prefer an incorrect but plausible-sounding statement to admitting that it does not know the answer, unless it has been specifically trained to do so through safety filters.20 For professionals in the Benelux who use AI for legal, medical or technical analysis, this presents an enormous risk. AI is not a source of truth, but a reflection of human language patterns optimised for a pleasant user experience.16
The search engine myth and the reality of token prediction
Many users in the Benelux approach ChatGPT as an evolved version of Google. This is a technological misconception. A search engine builds an index of the web and tries to show users the most relevant pages based on their search terms.12 Generative AI, by contrast, has no “knowledge” in the traditional sense of the word. The model is a vast collection of statistical weights in a neural network, designed to predict the probability of the next “token” (a word or part of a word) in a sequence.16
When a user asks a question, AI does not consult a database of facts. Instead, it “calculates” which words most logically follow the question, based on the patterns it encountered during training on billions of web pages.17 This explains why AI sometimes struggles with simple logic or arithmetic, yet excels at writing poetry or marketing copy. It is a creative system, not a factual one.14
Although newer versions use techniques such as “Retrieval-Augmented Generation” (RAG), in which AI first performs a search and uses the results to formulate an answer, statistical prediction remains at the heart of the process.12 The danger is that users confuse the fluency of the text with the accuracy of the information. This lack of “explainability” (the inability to determine exactly how AI arrived at a particular answer) makes it an unreliable tool for tasks where precision is essential.15
The hard truth about data storage and retention
One of the most persistent misconceptions is that data “disappears” after a session or that a paid subscription provides complete protection against data storage. However, the privacy policies of OpenAI and Microsoft are very clear on this: data is collected and retained.7
Data collection by OpenAI
As soon as an account is created, OpenAI collects account details such as names, contact information and payment details.7 More importantly, it collects “User Content”: all input (prompts), uploaded files and feedback provided by the user.7 Even for users with a paid account (ChatGPT Plus), OpenAI can use the data to improve its models by default, unless the user explicitly opts out through the privacy settings.5
Even when a user disables the “Chat History & Training” option, conversations are still retained on OpenAI’s servers for 30 days.25 This is done for “abuse monitoring”, checking that AI is not being used for illegal or harmful purposes.25 During this period, authorised employees or specialist third parties may access the data under strict conditions to investigate suspected abuse.26 This means there is never “zero retention” for the average user, regardless of whether they pay.
The hierarchy of data protection in AI services
| Service | Use for training | Data retention | Level of protection |
|---|---|---|---|
| ChatGPT (Free/Plus) | Yes (unless opted out) ^5 | Indefinite/30 days with opt-out ^25 | Consumer standard |
| ChatGPT Enterprise | No (off by default) ^5 | According to company policy (often 30 days) ^25 | Enterprise standard |
| Microsoft Copilot (Consumer) | Yes (possibly) ^28 | Variable, according to privacy policy ^28 | Consumer standard |
| Microsoft Copilot (Entra ID) | No ^29 | Within the tenant boundary (30-day logs) ^30 | Enterprise standard (EDP) |
| Azure OpenAI API | No ^31 | 30 days (0-day option on request) ^26 | Maximum control |
For businesses in the Benelux operating under strict regulations such as the GDPR, it is crucial to recognise that using the consumer versions of these tools almost by definition leads to a data security breach as soon as personal data is entered.8
The environmental footprint: an invisible cost
Beyond privacy and security risks, generative AI has a significant impact on the physical world. Training and running these models requires enormous amounts of energy and water, often out of the end user’s sight.
Electricity consumption at scale
Generating an answer with an AI model uses many times more energy than a traditional search query. Estimates indicate that AI-driven data centres will consume approximately 1,000 to 1,300 terawatt-hours (TWh) of electricity by 2030, double the current level.33 A single ChatGPT prompt uses an average of approximately 0.34 watt-hours of electricity.34 Although this seems negligible per query (comparable to running a television for nine seconds), billions of daily interactions add up to consumption comparable to that of tens of thousands of households.35
AI’s thirst: water consumption for cooling
Data centres produce enormous amounts of heat that must be removed to keep servers running. Many facilities use evaporative cooling, which literally consumes water.37 Researchers have calculated that a conversation of 20 to 50 questions with a model such as ChatGPT “drinks” approximately 500 millilitres of water.37 In 2023 alone, Google’s data centres consumed more than 29 billion litres of fresh water.38
The paradox is that this technology is often presented as “cloud-based” and therefore immaterial, while in reality it depends on scarce natural resources.37 For businesses that prioritise sustainability (ESG criteria), the uncontrolled use of energy-intensive AI models is a factor that must be included in a comprehensive cost-benefit analysis.
Legal consequences in the Benelux: the GDPR and the AI Act
Businesses in the Benelux operate within one of the world’s strictest legal frameworks for data and AI protection. Employees sharing data without a second thought is no longer merely an internal policy issue, but a potential legal minefield.
Warnings from regulators
The Dutch Data Protection Authority (AP) and the Belgian Data Protection Authority (GBA) have taken clear positions. Sharing personal data with AI chatbots without appropriate safeguards is classified as a data breach.8 The AP reported incidents in which doctors entered patients’ medical data into chatbots, directly violating professional confidentiality and privacy law.8 In Belgium, the Flemish Supervisory Commission (VTC) stresses that organisations must not transfer personal data to tools such as Copilot without first conducting a Data Protection Impact Assessment (DPIA).41
The impact of the European AI Act
The European AI Act has been in force since February 2025 and takes a risk-based approach to AI systems.42 Chatbots are generally classified as “limited-risk” systems, meaning they must meet strict transparency requirements: users must know they are speaking to a machine.43 However, once AI is used for tasks such as staff selection, creditworthiness assessments or education, the system may be classified as “high-risk”, requiring extensive documentation, human oversight and risk management.42
Businesses that tolerate “Shadow AI” (employees using unauthorised AI tools) expose themselves to fines of up to €35 million or 7% of global annual turnover.42 Legal responsibility rests with the employer, who, as the “data controller”, is responsible for employees’ actions even when they act against company policy.32
Data exfiltration case studies: Samsung and Amazon
The risks are not merely theoretical. Recent incidents at global technology giants show how vulnerable intellectual property is in the age of generative AI.
The Samsung incident: a lesson in source code security
In March 2023, Samsung discovered that employees in its semiconductor division had repeatedly shared sensitive information with ChatGPT.46 There were three specific incidents:
- A technician entered faulty source code from a database measurement program to find a solution.46
- An employee pasted code for identifying defective equipment into the chatbot for optimisation.6
- An employee uploaded a recording of an internal meeting to an AI tool to generate minutes.6
All this data became part of OpenAI’s training set, effectively exposing Samsung’s intellectual property.4 Samsung responded by immediately banning the use of generative AI for all its employees.4
Amazon and the risk to strategic insights
Amazon also warned its employees after ChatGPT responses were found to look suspiciously similar to internal strategy documents.46 Employees used the tool to create PowerPoint slides or polish text without realising they were feeding the company’s competitive advantages to an external party.6 Estimated losses from leaks of this kind can run into the millions, not only through fines but, above all, through the loss of strategic advantage.48
“Shadow AI” and the governance challenge
The term “Shadow AI” describes employees using AI systems without the knowledge or approval of the IT or compliance department.32 This is the modern equivalent of “Shadow IT”, but with a much greater impact because AI systems do not merely store data: they also process it and learn from it.49
Research by the National Security Alliance shows that 38% of employees admit to entering sensitive company data into AI tools without permission.45 In the Benelux, this problem is exacerbated by the high level of digitalisation and the pressure to be more productive. Employees see AI as a useful tool and bypass firewalls through their own mobile devices or personal accounts.50
Risks of uncontrolled AI use
- Exposure of trade secrets: Confidential information ends up in public datasets.32
- Loss of quality: Unchecked AI output (hallucinations) is communicated to customers.20
- Copyright infringement: AI generates content that infringes third-party copyright, for which the business may be held liable.20
- Bias and discrimination: Unknowingly using biased AI models in HR processes can lead to legal claims for discrimination.42
Strategic recommendations for businesses in the Benelux
To benefit from AI without compromising security, organisations must shift from a prohibition mindset to a governance mindset. Simply blocking ChatGPT is often counterproductive; it does not provide a safe alternative.6
- Develop an explicit AI policy
Businesses should create a document clearly stating which AI tools are allowed, who may use them and for what purposes.42 This policy must specifically define which data is “off-limits” (for example, customers’ personal data, annual financial figures that have not yet been made public, and source code).20 - Invest in AI literacy training
Employees need to understand that public AI tools are not personal friends and that their data is stored.20 Training should address the risks of hallucinations, how privacy settings work and the importance of human oversight of AI output.32 - Implement local, data-secure solutions
For organisations that want to use AI seriously, public consumer accounts are inadequate. The future lies in agentic AI systems that run locally. We build AI solutions that are trained and run locally. This keeps data within the customer’s own security boundary and prevents it from being used to train public models, a crucial safeguard for data security. Solutions such as Azure OpenAI provide a comparable isolated environment within a company’s cloud infrastructure,31 but fully local, customer-focused agentic models are preferred.29 - Use technical safeguards
Technologies such as “Data Loss Prevention” (DLP) can be used to prevent employees from copying sensitive patterns (such as Dutch citizen service numbers, or BSN numbers, and credit card details) to AI endpoints.50 Filters can also be configured to validate AI system output before it is sent to customers.41
Conclusion
Our appearance in De Tijd and the subsequent conversations across the region have taught us that the appetite for AI-driven productivity often exceeds awareness of the risks. ChatGPT is not a discreet assistant, but a commercial product designed to recognise patterns and to please.
In a world where data is the most valuable asset, thoughtlessly copying and pasting business information into a public chatbot is like leaving the keys to your safe in a busy marketplace. We therefore advocate a conscious approach to AI: use secure business environments, train employees in AI literacy and recognise that, regardless of your subscription, discretion in public AI tools is an illusion.
The illusion of privacy in paid accounts and the misconception that AI is a factual source of information pose a direct threat to the integrity of businesses. In a world where data is the new gold, thoughtlessly copying and pasting information into public AI is like leaving the keys to a safe in a busy marketplace. Only through a combination of strict governance, in-depth education and secure enterprise architectures can organisations in the Benelux harness the power of AI safely and responsibly. Responsibility for this transformation lies with senior leadership: privacy and security in the age of AI are not IT problems, but fundamental strategic priorities.
Sources
- Effect of anthropomorphism and perceived intelligence in chatbot avatars of visual design on user experience
- When Chatbots Feel Human: How Anthropomorphism Shapes Consumer Satisfaction, Trust, and Loyalty in AI-Driven Brand
- Effect of anthropomorphism and perceived intelligence in chatbot avatars of visual design on user experience – Frontiers
- Samsung Cold Case and the Shadow AI Flaw – Cyber Grant Blog
- How your data is used to improve model performance | OpenAI Help
- Samsung Engineers Feed Sensitive Data to ChatGPT, Sparking Internal Warnings
- Privacy Policy – OpenAI
- AI Chatbots Bring a Big Data Privacy Risk – DPO Consultancy
- Reward Shaping to Mitigate Reward Hacking in RLHF – arXiv (PDF)
- Reward Shaping to Mitigate Reward Hacking in RLHF – arXiv (HTML v5)
- Reward Shaping to Mitigate Reward Hacking in RLHF – arXiv (HTML v3)
- AI Engine vs Search Engine: Which One Should You Use? | WillDom
- When Google Falls Short: Smarter Research for CRE Pros Using Generative AI | Biscred
- Exploring the Differences: Search vs Generative AI | ClearPeople
- Predictive AI vs Generative AI – Red Hat
- What Are Large Language Models (LLMs)? – IBM
- LLM SEO: How to Rank Your Website in AI Search Engines – Tenet
- RLHF-V: Towards Trustworthy MLLMs via Fine-grained Correctional Human Feedback – arXiv
- Aligning Large Multimodal Models with Factually Augmented RLHF – arXiv
- Using AI Like ChatGPT at Work? Points to Consider for Safe Use – Legal News
- LLMs vs Search Engines: Which Traffic Actually Converts? – Bruce Clay
- How LLMs Access Real-Time Data from the Web – ML6
- Generative AI vs Predictive AI: What’s the Difference? – IBM
- Privacy Policy | OpenAI (UK)
- OpenAI ChatGPT Privacy Policy Requirements – Usercentrics
- ChatGPT and Privacy: How Closely Do OpenAI and Microsoft Take It? – Sulzer GmbH
- ChatGPT Pricing – OpenAI
- Does Microsoft Copilot Store Your Data? – Nightfall AI
- Difference in Data Privacy Between Copilot Free and Copilot Pro – Microsoft Learn
- Enterprise Data Protection in Microsoft 365 Copilot – Microsoft
- Azure OpenAI vs Public ChatGPT: Security and Privacy Considerations – Medium
- Shadow AI in the EU & Estonia: Legal Risks and Compliance – Magnusson
- Generative AI Power Consumption and Sustainable Data Centers – Deloitte
- We Finally Know How Much Energy and Water a ChatGPT Query Uses – Reddit
- How Hungry Is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference – arXiv
- The Real Environmental Footprint of Generative AI: What 2025 Data Tell Us
- From Cloud to Cup: How Much Water Does Your ChatGPT Drink? – IE Insights
- The Often Overlooked Water Footprint of AI Models – Julia Barnett
- Data Drain: Land and Water Impacts of the AI Boom – Lincoln Institute
- Dutch Data Protection Authority Warns About AI Chatbots and Personal Data Breaches
- AI and Personal Data – Flemish Government
- What Does the AI Act Mean for Employers and Employees? – Securex
- Using AI in Your Business – FPS Economy
- Artificial Intelligence Systems and the GDPR – Data Protection Authority
- What Is Shadow AI and How Can It Be Detected? – 2B Advice
- A Case Study on Samsung’s ChatGPT Incident – HumanFirewall
- Incident 768: ChatGPT Implicated in Samsung Data Leak – Incident Database
- 8 Real World Incidents Related to AI – Prompt Security
- Why Shadow AI Is the Biggest Compliance Threat – Staple AI
- AI Usage by Employees: Policy and GDPR Compliance – Reddit IT Managers
- AI Permissible Use Policy – GDPRWise
- Using AI in Your Dutch Business: GDPR and Compliance Risks – Law & More
- At the Crossroads of AI: ChatGPT OpenAI or Azure OpenAI? – Intwo
- AI Chatbots and AI Agents: Legal Concerns Under Belgian Law – ICT
