Balancing AI and human insights in scientific discovery: Challenges and guidelines

COMMENTARY Open Access Download: PDF

Recent advances in large language models (LLMs) have enabled machines to integrate web search, code execution, data analysis, decision-making, and even laboratory experimentation, as done in chemical discovery using the “co-scientist.”1 This artificial-intelligence (AI)-driven platform represents a pivotal moment in the evolution of systems. By using LLMs such as GPT-4 and Claude, co-scientists can autonomously design, plan, and execute complex chemical experiments based on simple natural-language prompts. Their capability lies in the ability to interpret plain-language requests, perform extensive data searches, synthesize information, and autonomously operate laboratory equipment through robotic application programming interfaces (APIs).


These capabilities can offer significant benefits for scientific writing and research proposal generation: for instance, they can assist researchers in articulating ideas more effectively, reducing language barriers, and streamlining administrative aspects of proposal preparation. In this context, the integration of AI into scientific proposal writing has accelerated in recent months. In particular, Google’s recent launch of the AI co-scientist, a virtual collaborator to assist scientists based on Gemini 2.0, has the potential to revolutionize the generation of novel hypotheses and research plans.2 Early collaborations with institutions such as Stanford University and Imperial College London have demonstrated its ability to independently hypothesize novel gene transfer mechanisms and suggest potential treatments for diseases such as liver fibrosis. Similarly, platforms such as Future House (https://www.futurehouse.org/) are emerging to further automate the scientific discovery processes, indicating a broader trend toward AI-driven research methodologies.


In fact, the adoption of LLMs in research is accelerating rapidly. A recent Nature survey reported that 81% of researchers have used AI tools like ChatGPT in their work.3 This trend highlights the urgent need to critically assess how these tools influence scientific workflows. This assessment is not occurring at the rate of change being produced.


The integration of AI into all stages of discovery, from hypothesis generation to interpretation of results, also raises significant risks and ethical concerns, including funding allocation. For example, the efficiency of AI-driven research may deprioritize projects that rely on intuition, reflection, and longer timelines, potentially stifling novelty and ambition.


AI use also opens avenues for misconduct, such as data fabrication or biased designs, without proper oversight. Moreover, it can manipulate public opinion and scientific narratives. A historical parallel is how the sugar industry shifted blame for health issues to dietary fat through manipulated research and lobbying, with long-term consequences in the United States. With LLMs able to generate persuasive content at scale, similar tactics could now be deployed more efficiently. In addition, control of advanced tools by a few companies or countries could monopolize discovery, using it as a means of dominance.


Another important point is lateral thinking. LLM-generated ideas show some creativity, complementing human contributions. Yet, the proportion of breakthrough discoveries has declined: a 2023 study analyzing 45 million papers and 3.9 million patents found a marked drop in the “disruption index” across disciplines since the 1940s.4 This is partly due to incentives favoring incremental projects. If LLMs are widely adopted for proposal writing, they may reinforce this trend by prioritizing ideas aligned with existing literature, limiting cross-disciplinary innovation. Trained on past research, they reflect historical biases and conventional paradigms, restricting radically new concepts. They are designed to fill gaps between known facts, not to generate truly novel insights.


To draw an analogy: imagine the time just after Newton had developed calculus and mechanics. AI could refine and scale this framework but would lack the insight of an Einstein or a Planck, capable of revolutionizing physics through relativity and quantum theory. AI might push classical mechanics to its limit, yet it would never confront questions such as whether Schrödinger’s cat is alive or dead.




Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(1444) Cited by(0)

Relative Articles