Skip to main content
AI offers powerful new capabilities for orthopaedic research, but meaningful insights still depend on expert human oversight and judgment.

AAOS Now

Published 8/25/2026
|
Sarah E. Lindsay, MD

Orthopaedic fundamentals: Artificial intelligence is an assistant, not a replacement, in orthopaedic research

At a Glance

  • AI can improve research efficiency, data analysis, and idea generation, but it remains dependent on human oversight.
  • Researchers must recognize AI limitations, including lack of transparency, hallucinations, and privacy concerns.
  • Human-in-the-loop design helps ensure AI-assisted orthopaedic research remains accurate, reproducible, and clinically meaningful.

Estimated read time: 4 minutes

Editor’s note: This article is part of “Orthopaedic Fundamentals,” a recurring AAOS Now series highlighting core principles widely embraced across the orthopaedic community. The series highlights ideas that can improve efficiency within and outside the OR, informing the training of surgeons as well as providing practical refreshers for experienced surgeons.

Artificial intelligence (AI) has boomed in recent years, and its growth is, at once, exciting and full of possibilities while also disconcerting. Although AI is a tool that has incredible power to supercharge research within orthopaedics, it also has many shortcomings and should be approached cautiously and thoughtfully.

AI offers powerful new capabilities for orthopaedic research, but meaningful insights still depend on expert human oversight and judgment.
Figure 1: AI is an assistant, not a replacement in orthopaedic research. Artificial intelligence can support elements of the research process, including data analysis, but human expertise remains essential for developing research questions, validating results, and interpreting findings in a clinical context.

AI is an intentionally broad term, and there are many misconceptions about what AI is and what it can and cannot do. AI is a technology that can simulate human learning and problem solving. It is an umbrella term that encompasses many different technologies, including natural language processing (NLP), machine learning (ML), deep learning, and generative AI.

---

Related Content

---

NLP enables computers to process and interpret human language. ML facilitates computer pattern recognition of data to make predictions or decisions about future data. Deep learning is a subtype of ML in which computers rely on neural networks (intricate relationships between data points) to evolve in real time with new data inputs to analyze complex data patterns. Generative AI has gained the most publicity and recognition in recent years and describes AI capable of producing new content based on user input.

AI-driven tools have significantly increased the efficiency and output of medical researchers through idea generation, complex model development, and enhanced data analytics. The use of generative AI has raised several concerns within the scientific community regarding the authenticity and accuracy of scientific work. While AI can convincingly and confidently generate research ideas, write code, analyze data, conduct literature reviews, and even write manuscripts, the accuracy of AI in performing these tasks can be highly circumspect. Ultimately, AI is a tool, not a panacea. It is only as good as the humans behind it.

Human-driven, AI-assisted research can be powerful. “De Novo Natural Language Processing Algorithm Accurately Identifies Myxofibrosarcoma from Pathology Reports” is one example of why it is important to have a so-called human in the loop (HITL). The objective of this study was to mitigate the primary issue present in many sarcoma studies: low numbers due to disease rarity. The study sought to build a new tool to facilitate the creation of large diagnosis-specific databases to answer targeted research questions.

The resulting NLP algorithm was designed to review the text of pathology reports to identify and confirm given sarcoma diagnoses for inclusion within the database. In the several years it took to develop the algorithm, the capabilities of AI expanded exponentially. In creating the algorithm, numerous shortcomings of emerging AI technology became apparent; through troubleshooting, solutions were found to address these issues. The lesson was straightforward: AI can accelerate research, but it cannot replace careful human judgment.

The black box problem
One of the key limitations encountered with AI is the black box problem. This describes the lack of transparency that is often inherent in AI-generated models, making it difficult to understand how those models arrived at their decisions. Even when a given model produces accurate results, the reasoning behind those results can be obscured within internal parameters. Like a middle school math test, the right answer is not just a number but also the ability to “show your work.”

Some AI models are nearly impossible to replicate because the steps in achieving the model outcome remain unknown and irreproducible. In the sarcoma NLP model, each step was human-designed and therefore clearly visualized, so it was obvious at each step of the algorithm why each decision was made. While this significantly increases labor on the part of those designing the algorithm, it ensures the steps are defined and the process is clear.

Beware AI hallucinations
Another limitation of AI is the hallucination problem. AI can generate factually incorrect information that sounds plausible and, to the uninformed consumer, may be mistakenly interpreted as correct. For example, AI may fabricate data, cite a scientific study that does not exist, or give the wrong explanation for something, all while sounding confident and authoritative.

The sarcoma NLP model was based on a series of predefined rules established by three sarcoma experts. For this reason, this model could not offer hallucinations, as it was tasked only with interpreting pathology reports based on rules given to it. This again puts the responsibility of rule generation and manual review on the algorithm designer.

Do not ignore privacy and security risks 
Finally, AI can pose significant privacy and security risks. Generative AI can require the sharing of protected health information (PHI) with public servers. As these models are constantly learning, this PHI can inadvertently appear during subsequent iterations. Users of many public AI tools sign end-user agreements that enable inputs to be used for model training and allow for the collection of personal information.

Furthermore, the copyright of AI-generated outputs remains nebulous. The advantage of models like the sarcoma model is that they can be applied within closed-loop systems, such as within the electronic health records (EHRs) of any given institution. External validation of the algorithm is possible because the algorithm itself consists of sequential steps and does not contain PHI.

AI affords orthopaedic researchers tremendous advantages, but it should be approached as an assistant and not as a replacement. It can help improve efficiency with repetitive tasks, particularly when working with big data. Whenever AI is used, the steps and methodology should be clearly visible and reproducible. Every step should involve some element of human review (HITL design).

As the capacities of AI continue to expand, it is necessary for orthopaedic surgeons to pause and consider that just because AI can technically do something does not mean it necessarily should be used for that purpose. There should be extra layers of caution in applying AI to medicine and medical research, because the outputs that are generated have implications for decisions regarding human health and lives. Good health policy decisions are contingent on sound clinical research. The onus is on researchers to ensure research outputs are accurate on every level (Figure 1).

Moving into this brave new world of rapidly evolving AI requires neither embracing nor rejecting these powerful technologies but rather walking that ever-elusive middle ground of critically considering and judiciously implementing new tools as they become available. By these means, AI can augment orthopaedic research rather than become a Trojan horse.

Sarah Lindsay, MD, is a current clinical fellow in musculoskeletal oncology at Mount Sinai Hospital of the University of Toronto in Toronto, Ontario. She completed her undergraduate and medical degrees at Stanford University in Stanford, California. She completed her residency in orthopaedic surgery at Oregon Health & Science University (OHSU) in Portland, Oregon. Upon completion of her fellowship in the summer of 2026, she will join the Department of Orthopaedics and Rehabilitation at OHSU as an orthopaedic oncologist.

References

  1. Wu S, Miao Y, Mei J, Xiong S. The rise of artificial intelligence in orthopedics: a bibliometric and visualization analysis. J. Multidiscip. Healthc. 2025;18:6037-6050.
  2. Chopra H. Annu, Shin DK, et al. Revolutionizing clinical trials: the role of AI in accelerating medical breakthroughs. Int. J. Surg. Lond. Engl. 2023;109:4211-4220.
  3. Faiyazuddin M. Rahman SJQ, Anand G, et al. The impact of artificial intelligence on healthcare: a comprehensive review of advancements in diagnostics, treatment, and operational efficiency. Health Sci. Rep. 2025;8:e70312.
  4. Prillaman M. Is ChatGPT making scientists hyper-productive? The highs and lows of using AI. Nature. 2024;627:16-17.
  5. Lindsay SE, Madison CJ, Ramsey DC, Doung YC, Gundle KR. De novo natural language processing algorithm accurately identifies myxofibrosarcoma from pathology reports. Clin. Orthop. 2025;483:80-87.