The Algorithm Doesn't Always Tell You What's True
Driving through Utah recently, I found myself between two large wildfires.
Smoke from the Cottonwood and Ironwood fires filled the sky. Winds were blowing 40 to 70 miles per hour. If you've never experienced something like that, it's difficult to describe how unsettling it is. At moments, looking at the smoke and knowing how quickly fire could move in those winds, it felt almost apocalyptic.
Later that day, I was back in town where the air was clear. I watched a Facebook video of the governor announcing a statewide ban on fireworks for the Fourth of July.
What surprised me wasn't the announcement.
It was the comments.
A surprising number weren't about fireworks or fires. People were making comments and jokes about AI data centers using water. The comments Facebook was showing me weren't a random collection of comments about data centers. They were predominantly critical of data centers.
I'm not opposed to data centers. I don't participate in anti-data-center groups or protests. So this didn't appear to be the case we often hear about where an algorithm learns your beliefs and shows you more people who agree with you.
It made me wonder about another possibility.
Was the platform using a major event—the wildfires—as an opportunity to give greater visibility to criticism of data centers?
We see a version of this in traditional media. A hurricane, wildfire, flood, or heat wave occurs, and the event becomes an opportunity to discuss a larger issue such as climate change. Sometimes that connection is scientifically relevant. There's still an editorial decision being made: this event will be presented through a larger frame.
Could a social-media platform do something similar, by deciding which comments and posts receive the greatest visibility?
Seeing one Facebook comment section doesn't establish that Facebook was intentionally promoting opposition to data centers. There are other explanations. Those comments might simply have generated higher engagement (a feedback response).
I wasn't seeing a representative sample of what everyone thought about the fires. I was seeing what an algorithm selected for me to see.
If platforms can decide which subjects receive visibility and which framing of those subjects receives visibility, then algorithms can influence the question we encounter before we start looking for answers.
I wasn't simply choosing which comments to read.
An algorithm was choosing which comments I saw.
That means the first question may be shaped before we even realize we're asking it. I thought the answer was simple. Data-center water use had become a popular topic, so people were bringing it into unrelated conversations.
I started thinking about something more interesting.
What if the problem begins with the first question we're given?
The Algorithm Doesn't Have to Tell You What's True
A recommendation algorithm doesn't necessarily need to decide whether something is true or false. Its job may simply be to determine what you're most likely to watch, read, click on, or engage with next.
Suppose you watch a video titled:
AI Data Centers Are Draining Our Water Supply.
The platform learns something from your behavior. You're interested in AI. You're interested in data centers. You're interested in water consumption.
What should it recommend next?
Probably another story about AI and water. Then another about data centers consuming water. Then perhaps one about a community concerned about a proposed data center.
None of those stories has to be false.
The problem isn't misinformation. The information can be accurate. The problem is that each new piece of information begins within the same frame as the first one. The recommendation system isn't likely to ask: What information would help this person understand the issue?
The objective may be: What is this person most likely to engage with next?
Those lead to different places.
There's Another Layer
Recommendation systems aren't the only thing determining what we see.
Platforms may decide which posts and comments receive greater visibility and which receive less. Moderation policies, ranking systems, engagement signals and internal platform decisions can all affect what appears prominently in a feed or comment section.
I've previously spoken on my podcast with a former Facebook employee and whistleblower about how platforms can amplify or throttle content for various reasons.
There are at least two forces shaping what we see:
What we choose to engage with, and what the platform chooses to show us.
Neither asks whether the information being presented is what we need to see to understand the larger problem.
I Found Myself Doing Something Different
I've spent considerable time investigating claims about AI data centers and water.
I could have started with:
Do AI data centers use too much water?
As I investigated, my questions changed to:
How is data-center water consumption actually calculated? Does a reported number represent water used directly for cooling, or does it include water associated with producing electricity? If electricity-related water is included, are we applying the same accounting method when discussing other technologies? Should water use be measured per unit of electricity consumed? Per unit of computation? Per unit of economic output? How much water would we attribute to charging an electric vehicle if we used the same methodology? If someone proposes moving data centers into space to eliminate terrestrial cooling-water consumption, are we comparing land-based and orbital data centers using the same system boundaries?
Those aren't additional answers to the original question.
They're different.
Different questions produce a different understanding of the problem.
Music Led Me to the Same Problem
I encountered something similar while investigating AI-generated music.A common question, often framed as a conclusion, is: Is AI stealing musicians' music?
If that's where you begin, you'll find an enormous amount of artists making the claim and information addressing it.
I found myself asking something different:
Did they, and if so, how did the recordings actually get into an AI training system?
Instead of beginning with stealing, I'm asking about the mechanism - how? Where did the audio files come from? Who obtained them? How were they stored? Who provided them to the training system? What permissions or licenses existed? The point isn't that the original concern is wrong. It's that the original question may already contain part of the conclusion we're investigating.
You Can Become Better Informed—and Miss the Problem
Imagine someone genuinely wants to understand an issue. They aren't trying to confirm a political belief. They aren't deliberately seeking misinformation. They're sincerely curious.
They watch one video.
They search for more information.
The algorithm sees their interest, and recommends another relevant video.
They read articles. They watch interviews. They follow people who discuss the subject. Eventually, they may know considerably more about it than they did when they started.
In one sense, they've become better informed. They've also become extraordinarily well informed inside the framework established by the first question. The more information they consume, the more confident they may become in that framework as seemingly independent pieces of information continue reinforcing it.
The algorithm isn't deceiving them. The articles don't necessarily need to contain false information. The experts don't necessarily need to be wrong. Everyone can provide useful information, while the investigation itself remains trapped inside the wrong question.
The Feed Follows the Question
We tend to think the solution to poor information is more information.
Sometimes it is.
Sometimes, consuming more information within incomplete framing may deepen the problem and lead to an inaccurate conclusion.
Suppose I begin by asking:
Why are data centers using so much water?
I've already embedded an assumption with that question: data centers are using so much water.
Now I search.
The search results address data-center water consumption. Videos explain cooling systems. News stories cover communities concerned about water. Social-media algorithms recommend similar content. Everything appears to confirm that I'm investigating the right problem.
What happens if I ask:
How is data-center water use calculated?
I encounter information about direct versus indirect water use, cooling technologies, electricity generation, geographic differences, system boundaries and competing accounting methods.
I haven't rejected the concern. I've changed the investigation.
The same applies to AI music.
Instead of: Why is AI stealing musicians' work?
Try asking, How does music get into an AI training dataset?
The second question doesn't predetermine the answer. It gives the investigation somewhere to go.
That investigation has led me somewhere unexpected: I still haven't found a precise answer. It's possible that some of the music being attributed to “AI stealing” wasn't copied from an original recording, rather it was recreated based on information about an artist—their sound, style, metadata, and the kinds of music they typically release.
That's a possibility I'm still investigating, not a conclusion. I've spoken with A&R record executives about this, and haven't found an explanation of how the recordings in question entered the training system.
That's the problem I'm interested in.
If “AI stole it” is both the conclusion and the question we start with, how do we investigate whether that's actually what happened?
Thinking Beyond the Feed
Algorithms are not entirely to blame. The larger problem may be misinformation. Misinformation doesn't necessarily have to be where the problem starts. It can be what allows a questionable premise to keep going.
Algorithms may reinforce our initial assumptions. News and editorial decisions can reinforce them. Social conversations can reinforce them. The pattern can begin before any of that.
It begins with the first question. Once we accept that question, we naturally start looking for answers to it.
Then the algorithm helps.
It gives us more.
More videos. More articles. More experts. More examples.
If the premise behind our original question is correct, that may help us understand the issue.
If the premise is wrong, and the information we're given continues to support it, something has to give. Information may be incomplete, taken out of context, misleading, or simply wrong. Misinformation may not have created the original question. It may be what allows us to keep answering it.
We can accumulate an impressive amount of knowledge without ever asking whether the investigation began in the right place. Thinking clearly sometimes requires doing something that recommendation systems aren't designed to do: Step outside the question.
Ask what assumptions it contains. How else the problem could be framed? What information you would need to discover that your initial understanding was wrong?
Most importantly:
Ask what question you would be investigating, if you had never encountered that first headline, video or post.
The more I investigate complex issues, the more I suspect that the quality of our conclusions depends less on how much information we consume, than on whether we're willing to question the first question.
Once the first question becomes accepted, we can spend enormous amounts of time finding increasingly sophisticated answers to it,
Without ever asking whether it was the right question in the first place.
About the Author
Daniel Stih (danielstih.com) is an aerospace engineer, software engineer, indoor environmental consultant, and author of 12 books. For more than 30 years, he has investigated complex problems spanning engineering, technology, the built environment, and human decision-making. His work explores how evidence, assumptions, and systems shape the conclusions we draw—and whether we're solving the right problem. Learn more about his approach in Why I Think This Way.