Ever notice how chatbots often insist on making up an answer instead of admitting defeat? Like a student desperately scribbling nonsense to avoid handing in a blank page, large language models (LLMs) seem to live by the make-it-up-as-you-go principle.
Many also stumble when bridging two areas of knowledge that don’t have an obvious link. They tend to give a superficial response instead of following the trail.
Sure, it’s annoying when a search snafu derails your goal of unlocking the flying car in Grand Theft Auto VI. But when the fumbled response involves synthesizing medical research findings, the stakes get a lot higher.
Given that PubMed lists over 40 million citations and abstracts, finding the right information in the vast database is a job that begs to be outsourced to AI. But more often than not, a query spanning more than one knowledge realm will result in answers that are outdated, incomplete or plain wrong.
“Since knowledge is so very fragmented, data is all fragmented,” said leader of the Neural Dynamics Lab Ayan Paul in an interview with Northeastern Global News. Paul is a research associate professor at the Institute for Experiential AI.
Headed by the lab’s machine learning engineer Nihar Sanda, a group of researchers found a way to teach AI to dig deeper.
In a paper that appeared in Bioinformatics in July, the team presented an algorithm called eGoT. Standing for “enhanced graph of thoughts,” it teaches LLMs to answer biomedical questions by pulling together evidence from different domains. An algorithm like eGoT works as a set of instructions for the computer system. It tells the machine how to perform a task — in this case, information synthesis and retrieval.