Are LLMs Good at Finding an Edge?
· 2 min readLLMs
At its heart, an LLM is a tool that processes language. It reads what you ask for and produces the text most likely to satisfy it. This simple LLM is very unlikely to succeed at finding an edge, because it can only produce words from itself. While it is trained on many texts, including financial texts, that training at best can explain a financial idea or trading strategy at a high level. So if you were to ask an LLM "what is pairs trading?", it would probably give an answer on par with a Wikipedia article. And this is probably the best an LLM can do on its own; it certainly will not be sufficient to find an edge.
But the technology that most LLM tools provide these days goes well beyond mere language manipulation. Several techniques, including Retrieval-Augmented Generation (RAG) pipelines allow for fetching of external data to inform the inputs and outputs of an LLM. Continuing with the idea of pairs trading, a RAG system built on financial documents and journalism would allow the LLM to understand how pairs of commodities might be related. It would not be sufficient to run a pairs trade on its own, but it would be sufficient to reject pairings that are bad because they totally lack a relationship with each other, or ones that are so commonly observed that any alpha has long since decayed. Being able to reject bad strategies early is still useful in finding an edge, in that resources spent designing backtests are not wasted on bad ideas. With the right tooling around it, you can achieve a lot more than what seems possible on paper with a frontier model.
In that sense, you can think of LLMs as another tool to find an edge: They can exponentially reduce the cost of finding alpha. You wouldn't really ask "Are IDEs Good at Finding an Edge?", when you think from that perspective. The answer depends on whether they fit the researcher's workflow, but broadly you can reasonably conclude using first principles that an IDE can save a lot of human hours. Fewer human hours to do some research, less compute used up to do research that could have easily been ruled out means money saved. We have put this to practice ourselves as quantitative researchers: we run multiple live strategies with a Sharpe ratio above 1.8.
There is, however, a long way to go before LLMs can independently do all the work themselves. Even a better-defined task, building an entire product from a single prompt, is still beyond them. While recent state-of-the-art models have gotten really good, we believe that we're still far away from completely independent research. For now, a human with tools, LLMs among them, is still the strongest setup.
A shovel has never dug a hole in history, only humans with shovels have. Happy Digging!
P.S. We truly believe the thesis of human + tools. This is why we built Caliper Trading. It's a harness we developed through extensive research and trading our own money, live.
This is not financial advice.
