NYU's Study: General AI Outperforms Specialized Medical Tools

Aug 10, 2026 · 4 min read

NYU's Study: General AI Outperforms Specialized Medical Tools

General-purpose AI models like ChatGPT, Claude, and Gemini outperform specialized medical tools in answering clinical questions, according to a study by NYU. Doctors in the study consistently ranked general AI models higher, challenging the assumption that specialized tools are inherently superior.

Source

Watch the Reel

Medical AI: Do Specialized Tools Really Outperform General-Purpose AI?

The assumption that specialized medical AI tools outperform general-purpose AI in clinical settings has been challenged by a recent study. Scientists at NYU conducted a rigorous test comparing a $3.5 billion medical AI, Open Evidence, and another widely used tool, UpToDate, against general-purpose AI models like ChatGPT, Claude, and Gemini. The results were surprising: the general-purpose AI models consistently outperformed the specialized medical tools.

Context: The Evolution of AI in Healthcare

AI has become integral to modern healthcare, assisting doctors in diagnosing, treating, and managing patient care. Specialized medical AI tools, such as Open Evidence and UpToDate, were designed with the explicit purpose of aiding healthcare professionals. Open Evidence, for instance, has raised $210 million and is already in use in many hospitals. UpToDate, another well-known tool, has been a staple for doctors for years. These tools are built on Retrieval-Augmented Generation (RAG), which pulls data from vast databases to provide answers. However, RAG can sometimes retrieve incorrect information, which can degrade the quality of the answers.

The Study: Comparing Specialized vs. General-Purpose AI

NYU scientists took 100 real clinical questions that doctors had inputted into AI systems and ran them through both specialized and general-purpose AI models. Twelve doctors then graded each answer blindly, meaning they did not know which AI model provided the answer. The general-purpose AI models—ChatGPT, Claude, and Gemini—consistently received higher rankings from the doctors.

Key Findings

  • Performance Metrics: General-purpose AI models outperformed specialized medical AI tools in every single question.
  • Blind Testing: All 12 doctors ranked the general-purpose AI models higher than the specialized tools.
  • Data Retrieval Methods: Medical tools like Open Evidence and UpToDate rely on RAG, which can sometimes retrieve incorrect information, leading to potentially flawed answers. In contrast, general-purpose AI models like Frontier AI have knowledge baked into their training, making them more reliable in some cases.

Why This Matters

The study highlights a critical issue in the healthcare industry: the assumption that specialized AI tools are superior to general-purpose AI models is not always accurate. This misconception can lead to the adoption of expensive, specialized tools that may not perform better than freely available alternatives.

Implication for Hospitals

Hospitals and healthcare providers are already making real decisions based on AI tools. The findings from the NYU study underscore the importance of independent performance testing before adopting any AI tool. It's not enough to rely on marketing claims or assumed superiority; actual performance testing is crucial.

Financial Considerations

The cost difference is significant. For example, a specialized medical AI tool might cost $700, while a general-purpose AI model could be free. Given that general-purpose AI models performed better in this study, hospitals could save millions by opting for more cost-effective solutions.

Practical Tips for Healthcare Providers

Conduct Independent Testing

Before adopting any AI tool, conduct thorough, independent performance testing. Compare it against other available tools, including general-purpose AI models, to ensure it meets your needs and performs as expected.

Consider Cost-Effectiveness

While specialized tools may offer unique features, consider whether the added cost is justified by improved performance. General-purpose AI models might offer similar or better performance at a lower cost.

Stay Updated with Research

Keep abreast of the latest research and studies in medical AI. The field is rapidly evolving, and new findings can significantly impact your choices.

Promote Transparency

Encourage transparency in AI tool evaluations. Ensure that the performance claims made by AI tool providers are backed by independent, verifiable data.

Important Takeaways

  • Specialized medical AI tools do not always outperform general-purpose AI models.
  • Independent performance testing is crucial before adopting any AI tool in healthcare.
  • The cost of specialized AI tools may not justify their performance advantages.
  • General-purpose AI models often have knowledge baked into their training, making them more reliable in some cases.

Conclusion

The NYU study serves as a wake-up call to the healthcare industry. It challenges the long-held belief that specialized medical AI tools are always superior to general-purpose AI models. By conducting rigorous performance tests and considering cost-effectiveness, healthcare providers can make more informed decisions about the AI tools they adopt. This approach ensures better patient outcomes and more efficient use of resources.

Summary

Key points

  • General-purpose AI models outperformed specialized medical AI tools in addressing 100 real clinical questions posed by doctors.
  • Doctors ranked the general-purpose AI models higher than the specialized tools in blind evaluations of the answers provided.
  • The study demonstrates that general-purpose AI models can provide more reliable answers for medical questions, unlike specialized tools which use Retrieval-Augmented Generation (RAG) and sometimes retrieve incorrect information.
  • The study challenges the common belief that specialized AI tools are always superior to general-purpose AI models in medical settings.
  • The findings suggest that hospitals should conduct independent performance tests before adopting any AI tool in healthcare
  • The cost of specialized medical AI tools can be much higher than generic equivalents
Answers

FAQ

The NYU study aimed to determine whether general-purpose AI models like ChatGPT, Claude, and Gemini could outperform specialized medical AI tools in addressing clinical questions. The study sought to challenge the prevalent assumption that specialized tools are always superior in medical settings.

Discussion

Comments

Be the first to comment.

Similar reads based on topic and creator.

Recent articles

Fresh deep dives from the latest Reels we unpacked.

View all