Probabilistic Automation Systems
Presentation given as introduction for a discussion on "AI". Some funny stuff to start with: "your business AI"
What "AI"?
I write "AI" between quote marks, because the term is poorly defined and can refer to very different systems. Coined in 1956 by John McCarthy it's now widely used for marketing purposes. Anything that uses an algorithm could be called "AI". Furthermore the term "AI" is misleading to what it is. There's no intelligence behind it. Let's start with some issues about that. Systems like chatGPT for example are just statistic models to predict which character, which word and which sentence is likely to come next based on huge piles of data. Reason for Emily Bender to call it an stochastic parrot.
It's a container for a wide variety of different systems:
- large language models,
- chatbots
- translation systems
- picture labeling systems
- pattern recognition systems
- picture generating systems
- ... (fill in the gap)
The term is obfuscating the way these models actually work:
- the output is defined by statistics invoked by an algorithm, there is no intelligence
- "intelligence" feels like an independent system, obfuscating the responsibility of the developers and the company
The term is anthropomorphising:
- misconception that systems have an “intent”
- misleading users expectations and perceptions
- because we humans inevitably associate communication with a persona
For further read and alternatives, see this paper about De-anthropomorphising “AI”.
What’s the problem?
The set up, maintenance and use of "AI" comes with serious problems. In the following paragraphs the main ones.
Energy and water
Energy consumption of data centers is surging and the investments in new data centers is exploding. The International Energy Agency published a report about it. In this report it says that the capacity of "AI" factories have more than tripled in capacity in the past 18 months, causing growth in energy demand. "The global electricity demand of data centres – the critical infrastructure for training and running AI models – grew by 17% in 2025, in line with IEA projections. Electricity consumption from AI-focused data centres grew even faster, surging 50% in 2025." The investment in new data centres is expected to grow form 400 billion USD in 2025 with 75% in 2026.
An example to feel the scale:
- Amsterdam households annually use 869 GWh
- the 2 biggest datacenters in Amsterdam annually use 807 and 779 GWh respectively
- one of them will use 1.095.000 $m^3$ cooling water annually
- 25.000 people in the Netherlands on average use 1.077.000 $m^3$
We are worldwide consuming more energy instead of less as we should and we already are struggling to keep up with the target to keep global worming within 1,5 ºC.
Exploitation
Probabilistic Automation Systems are more than just algorithms. They need human labour too. There's a hidden workforce to do jobs like:
- annotation
- labeling
- transcription
- moderation
Most of this work is concentrated in Africa and south & southeast Asia which often comes with lack of regulations, contracts, (mental) healthcare, etc. See for example the research by the Deutsche Gesellschaft für Internationale Zusammenarbeid Or have a look at these documentaries:
- How big AI companies exploit data workers in Kenya | DW Documentary
- Training AI takes heavy toll on Kenyans working for $2 an hour | 60 Minutes
There are some good initiatives too:
Bias
Probabilistic Automation Systems are not representing all available information equally and they certainly are not a representative reflection of the worlds population. They all suffer from:
- gender bias
- racial bias
- age bias
- political bias
- … (fill in the gaps)
The origin of bias is multifold:
- data sets are never neutral and are currently dominated by western sources
- algorithms make selections from these datasets and introduce bias by design
- moderation and labeling are human activities and introduce subjectivity from the moderators Example: Comprehensive Study on Bias in Artificial Intelligence Systems
The wide use of these systems are causing also information contamination:
- information now published has a growing synthetic origin
- systems will be trained on partly synthetic fabricated data (also causing more bias)
- recent research by scientific publisher Nature shows: already 0,5% of references in scientific papers do not exist
Decline in skills
Research suggests decline in cognitive skills after frequent use of LLM’s. Example: decline in skills of physicians through use of pattern recognition Or read this study: AI-overdependence and human cognitive decline
We see this also happening in our fablabs and during fabacademy. Students are more and more using LLMs to write code or even design electronics. So in the end it's questionable if and what they learned during the course.
IP rights
Probabilistic Automation systems need data to get trained. This goes for LLM's, labeling systems or any form of "AI". Most of this data is roamed from the internet. All information ever published could be fed into the data sets of Anthropic, OpenAI and others. For most of it no compensation was given to the owners of the IP rights. And bots don't bother about CC statements either. Even if you implement some metadata to prevent bot searches, it's not certain that they react accordingly.
Some big IP-owners arranged an agreement with one or more tech companies. The Guardian for example has a strategic partnership with OpenAI about "cooperation", leading of course to a reduction of journalists.
Make no mistake about stand-alone systems. They may seem ok from a personal or company privacy point of view, but they are nevertheless also trained on data from others, illegally obtained in my view.
For some more info see the World Intellectual Property Organization.
Privacy
Linked to IP-rights is the violation of privacy rights. Some ways in which this will occur:
- mining from the internet includes personal data (facebook, LinkedIn, etc.)
- more or less secretive use of data, like WeTransfer tried to incorporate in their policy (and had to come back to it)
- automatic speech transcription systems in e.g. Teams sends this data to microsoft
- even some medical facilities use automatic speech transcription to summarise patient consultations
- privacy breaches like the Claude.app that was forcing access to users' browsers
Here are some articles about the subject in MIT technology review, 18th of July:
- How LLMs could supercharge mass surveillance in the US How LLMs could supercharge mass surveillance in the US
- Is the Pentagon allowed to surveil Americans with AI?
- What AI “remembers” about you is privacy’s next frontier
- A major AI training data set contains millions of examples of personal data
Democracy
With "AI" we give big tech enormous power causing a threat to democracy:
- ChatGPT, Gemini, Claude, etc. try to get dominance with massive investments
- shake out of smaller competitors (compare the main operating systems in use) and concentration of power in the remainder
- no public liability or democratic control over "AI" technology
- most systems are non-transparent and have no possibility to check algorithms, data and their effects
- growing power and growing influence gives a growing risk of interference in democratic processes either deliberate or unintended
What the fab?
As a fab community we have the responsibility to reflect on these issues and discuss how we want to deal with it:
- what do we see as ethically acceptable use?
- what do we want our students to learn about and from such systems?
- how do we deal with students that use it and learn nothing?
- how does open source relates to the data theft by the big tech systems?
We had a start of this discussion during our event in Bottrop. We focussed at that time mainly on the teacher - student aspect, but we have to continue the discussion in a broader sense. The technology is not going away and we should actively make it work for people and not for big tech. Apart from doing fabacademy and fabricademy every year in Waag, we also have the mission to make technology accessible for a wider public. This includes showing and explaining how things work and what the benefits and down sides are. This applies also to what is called "artificial intelligence". So please continue this discussion in your own environment too.
Photo from from the audience


