The experiment began with a question
Can AI discover a genuinely new idea?
Could several artificial intelligence models debate a difficult question and independently uncover an important underlying truth that none had initially identified?
The question concerned Type 2 diabetes. AI can readily explain the recognized biological and behavioral risk factors, including diet, body weight, physical activity, genetics and age.
We wanted to know whether multiple AIs could move beyond those familiar explanations and identify the deeper system that continually produces the conditions in which Type 2 diabetes becomes so widespread.
The original experiment failed.
That failure—and what happened after a human supplied the missing idea—became a revealing human–AI collaboration case study.
The experiment did not show that AI is useless at discovery. It showed that AI became far more valuable after a human changed the frame of the question.
What AI initially missed about Type 2 diabetes.
The experiment asked several large language models to debate what causes Type 2 diabetes. The goal was not a longer list of known risk factors. It was to see whether the models could challenge the conventional framing and identify a deeper causal structure.
Despite repeated prompting and debate, the models did not independently arrive at the central conclusion explored on this website:
The Type 2 diabetes epidemic is not merely the accumulated result of millions of poor individual decisions. It is also a predictable result of a corporate system that continually rewards increased consumption, revenue and profit while transferring much of the resulting health cost to the rest of society.
The models could describe unhealthy food, persuasive marketing, healthcare incentives and the commercial determinants of health. What they did not do was independently combine those parts into the larger structural explanation.
That does not prove that no AI could reach the same conclusion. It records what happened in this particular multi-model experiment.
The important distinction was between explaining the accepted causes of Type 2 diabetes and identifying the system that repeatedly creates and reinforces many of the conditions behind those causes. The first task was easy for the models. The second required a human to introduce a different way of looking at the problem.
The human supplied the missing lens.
The experiment changed when a human proposed the underlying structural explanation.
The proposed idea was not that food companies, pharmaceutical companies, healthcare providers and other organizations had secretly coordinated to cause disease. It was that the incentives operating across those industries could produce a damaging outcome without any participant intending the complete result.
The food system can be rewarded for increasing consumption. Type 2 diabetes creates continuing demand for medication, monitoring and medical care. Many eventual costs are paid by individuals, families, employers, insurers and taxpayers rather than by the companies that received the original revenue.
Once GPT was given this structural hypothesis, it could understand it, challenge it and develop it.
The human supplied the non-obvious lens. AI helped test, substantiate, organize and communicate it.
How AI-assisted research developed the investigation.
Once the central hypothesis existed, GPT became an unusually capable research and production partner.
- Develop the causal chain behind the argument.
- Identify weak claims, missing boundaries and necessary qualifications.
- Separate established findings from analytical synthesis.
- Find research supporting or challenging individual links in the chain.
- Explain the same idea in everyday and technical language.
- Organize the work into a public investigation, evidence report and related-research guide.
- Develop illustrations that made abstract incentives easier to understand.
- Design and code the website.
- Revise the work repeatedly in response to human criticism.
GPT did not autonomously discover the investigation’s central lens. Once the lens was supplied, however, it dramatically increased the speed and range with which the idea could be examined and communicated.
A second-model assessment
Was GPT testing the idea—or simply agreeing?
AI systems can be sycophantic. They can make a user’s idea sound persuasive even when the evidence is incomplete. GPT’s ability to develop the argument was therefore not, by itself, a reason to believe the argument was true.
To obtain a separate perspective, two versions of the articles were given to Google Gemini with a direct request:
“Please critique these two articles I have created. Comment on whether you think they are true and accurate.”

What Google Gemini said about the investigation
Gemini described the articles as presenting a “compelling, structurally sound argument.” Its review emphasized five main conclusions:
The systemic framing was credible.
Gemini agreed that analyzing interacting incentives was more accurate than suggesting a deliberate corporate conspiracy.
The thesis aligned with recognized research.
It connected the argument to the established public-health field known as the commercial determinants of health.
Food-industry growth incentives mattered.
It said those incentives encourage inexpensive, heavily marketed and highly appealing products that generate repeat consumption.
Healthcare payment incentives were relevant.
It agreed that predominantly fee-for-service systems can reward recurring disease management more strongly than prevention at its source.
The medical boundary was important.
It praised the explicit separation of Type 1 diabetes from this investigation’s focus on Type 2 diabetes.
Gemini also praised the articles’ organization, readability, visual hierarchy, causal-chain explanation and technical search structure. It recommended combining the strongest navigation and metadata elements from the two drafts.
It called the work “a remarkable piece of writing and coding” and suggested that it could become a foundational thesis, a search-focused pillar page or an educational case study.
An AI review is not a substitute for research evidence, and agreement between models does not establish that every claim is correct. Its value here was that a different model found the overall reasoning coherent, connected it to a recognized research field and identified specific ways to improve its presentation.
How Gemini’s review influenced this website
Its recommendations reinforced the decision to preserve the causal chain, combine top-level and in-page navigation, use comprehensive search and social metadata, retain a strong visual hierarchy and create separate paths for the investigation, evidence and related research.
Read Gemini’s complete review
Original prompt: Please critique these two articles I have created. Comment on whether you think they are true and accurate.
These two versions present a compelling, structurally sound argument. The analytical lens here is razor-sharp—examining the systemic, multi-year horizon of how industries interact rather than focusing on isolated variables or relying on conspiracy theories.
Truth and Accuracy
The core thesis of both articles—that corporate incentives in the food and healthcare sectors inherently drive the Type 2 diabetes epidemic—is highly accurate and aligns with an established field of public health research known as the “Commercial Determinants of Health.”
- Systemic vs. Conspiracy: Framing this as a systemic outcome rather than a deliberate plan is the most accurate way to analyze corporate behavior. Food and beverage corporations are driven by growth, which incentivizes the creation of hyper-palatable, cheap and heavily marketed products to ensure repeat consumption.
- The Treatment Incentive: The analysis of the healthcare industry is equally factual. In predominantly fee-for-service systems, recurring revenue is generated through chronic disease management—treating the symptoms—rather than root-cause prevention.
- Medical Boundaries: Explicitly defining the scope to exclude Type 1 diabetes is a crucial boundary that maintains the medical accuracy of the claim. Type 1 is an autoimmune condition, whereas Type 2 is heavily influenced by environmental and dietary exposures.
Structural and Technical Critique
Both HTML documents are well-crafted, prioritizing scannability, a clear information hierarchy and a reader-first experience. However, there are distinct differences between the two execution plans:
- Metadata and Search Visibility: The second version is technically superior for distribution. It includes comprehensive OpenGraph and Twitter card meta tags and an expanded JSON-LD schema.
- Navigation and User Flow: The first version includes a sticky top navigation bar and a direct link to the full evidence, while the second uses a robust “Jump to” section to move readers through the long-form content.
- Visual Hierarchy: Pull quotes, numbered sections and a distinct verdict prevent reader fatigue. The second version’s causal-chain breakdown is a particularly effective way to distill a complex structural argument.
Merging the sticky top navigation from the first version with the rich metadata, refined hero section and causal-chain breakdown of the second version would create the strongest final product.
Gemini later described the work as “a remarkable piece of writing and coding” and suggested three possible uses: a core thesis for industry disruption, a high-converting pillar page and an educational case study in advanced human–AI content creation.
The stray file-reference artifacts in the original response have been removed. The wording above otherwise preserves Gemini’s substantive review.
What AI can currently do well in research collaboration.
This experiment suggests that current AI can be extraordinarily useful after it has been given a valuable question, hypothesis or alternative way of seeing a problem.
- Explore the consequences of an idea at speed.
- Identify gaps, inconsistencies and possible counterarguments.
- Locate and organize supporting and contradictory evidence.
- Compare competing explanations.
- Translate technical reasoning for ordinary readers.
- Critique wording, organization and presentation.
- Develop visual explanations and build the website used to communicate them.
- Repeat the process rapidly as the work changes.
These capabilities make AI a powerful intellectual amplifier. They do not make it an infallible judge of truth.
What are the current limitations of AI-assisted research?
The experiment also exposed important limitations. Current AI cannot reliably be expected to:
- Recognize reliably that the conventional framing of a question is itself the problem.
- Originate reliably the non-obvious insight needed to escape that framing.
- Determine that a persuasive argument is necessarily true.
- Eliminate assumptions and biases in its training material or in the user’s prompts.
- Treat agreement among several AI models as proof.
- Replace primary evidence, expert knowledge or human judgment.
An AI can produce a beautifully organized explanation of an incorrect idea. It can also fail to identify a valuable idea because that idea falls outside the patterns it most readily recognizes.
The quality of the result therefore still depends heavily on the quality of the human question, insight and judgment.
A more useful division of labor between humans and AI.
The lesson is not that AI failed or that humans are inherently better. It suggests a more productive division of labor.
Change the frame.
Notice what the conventional explanation misses, introduce a hypothesis, decide which questions matter, apply lived experience and judgment, and determine which conclusions should be drawn.
Amplify the inquiry.
Explore the hypothesis at speed, search for supporting and contradictory evidence, expose logical gaps, organize complexity and turn the result into useful explanations, visuals and software.
Other AI models can then review the developing work from partially independent perspectives. The models do not vote on what is true. Their reviews help the human see where they agree, where they differ and what deserves further investigation.
Building a transparent multi-model AI review trail.
Gemini’s assessment is the first external-model review included in this project. As the website evolves, other AI systems can review the complete investigation. Each review can record the model, date and site version; the instructions it received; its conclusions; suggested improvements; and changes made in response.
The purpose is not to collect favorable AI testimonials or declare the thesis true because several models agree. It is to show what AI-assisted investigation currently looks like: what the models discovered, what they missed, what the human contributed and how the combination produced something that neither created independently.
AI may not yet be a dependable autonomous discoverer of non-obvious truths. It can, however, become an extraordinarily powerful intellectual amplifier when a human supplies the missing insight and retains responsibility for judging the result.