Does Claude Code Investing Work? My Backtest results...
•By Dave Wang
A common question I get is if AI for investment research actually works.
That's something I wanted to test (with numbers) this week.
In February I shipped a deep-research dossier (using Claude Code) on Singapore’s Equity Market Development Program. (Link to the original)
A summary of the thesis I wanted to test: Singapore is committing S$6.5 billion to revitalize Singapore’s small and mid-cap segment.
The capital injection should (in theory) drive a flywheel of liquidity, institutional discovery, and re-rating in the Small/mid cap universe.
I scored about 75 Singapore-listed names through a 5-factor framework I designed (flow sensitivity, research consensus, fundamentals, index visibility, theme alignment), then handed the universe to Claude Code to flesh out the details.
The output was a top 15-name basket organized into three conviction tiers.
Nine weeks later, the question is simple: did the basket actually outperform the market?
Like any responsible investor, you want to backtest your results. So that's exactly what we are answering this week.
...and I'll also give you tips on how you can get the most out of using AI for your investing frameworks.
The Backtest
Here's the backtest methodology I ran.
Window: Feb 20 close (publication day) to Apr 24 close. We are comparing our basket of names vs the STI (Singapore stock index)
Headline: the 15-name basket beat the STI by 858 basis points.
Top 15 basket (equal-weighted): +6.7 percent
STI benchmark: -1.9 percent
Spread: +8.6 percentage points
The STI was actually negative over the window, which makes the absolute return on the basket the more honest read. The basket was up ~7% in 9 weeks, or ~40% annualized.
We also wanted to measure the hit rate (grouped by conviction tier):
Three names did most of the heavy lifting, all semiconductor-exposed: UMS Integration +58.5 percent, Frencken +40.7 percent, Valuetronics +28.0 percent.
That clustering is not a coincidence. The dossier flagged “AI supply chain / semiconductor equipment” as the strongest “New Singapore” theme inside the EQDP fund mandates, and the back-test confirmed it.
Two names beyond the semis also worked well: Sheng Siong +12.7 percent on consumer-staple defensiveness, and Centurion +5.1 percent on accommodation niche tailwinds.
The basket had two real misses, both Tier 2: Yangzijiang Financial dropped 22.9 percent as the market caught up to the opacity of its investment portfolio, and SIA Engineering fell 9.1 percent off its premium valuation.
Most of the rest clustered in a narrow band around flat to slightly negative, which means the framework largely sidestepped serious losses and stacked the wins.
So the basket’s batting average was 9 of 15 names directionally beat STI. Three of those wins were big enough to carry the rest.
That’s the headline grade. Now look at the underlying signal by tier.
What this tells me: the framework’s primary edge in this back-test was identifying the theme (AI semis as the EQDP-aligned winner) rather than the individual name ranking. The conviction tiers were directionally right at the basket level but noisy at the per-tier level over 9 weeks.
The Lesson
Two thoughts on what this back-test actually validates.
The steel man is the alpha
The 5-factor framework was not generic. It was specific to this setup:
Flow sensitivity to capture which names move most on incremental institutional buying.
Research consensus to triangulate which names the appointed fund managers are likely to gravitate to.
Liquidity-vs-market-cap mismatch to find names that trade more than their float would suggest, indicating active interest.
Theme alignment to map names to the explicit fund mandates (“New Singapore” technology, data, services).
Fundamentals to filter out value traps.
Those five lenses are what AI couldn’t have generated on its own. They came from understanding the macro structure of the EQDP, the fund manager ecosystem, and how SMID-cap re-rating tends to play out.
What AI did was apply the framework to 75+ names with consistent rigor and produce a 50,000-word dossier in hours. That used to be a six-week analyst project.
The pattern keeps repeating. The differentiated view comes from the human steel man, and the breadth comes from AI.
One point I always emphasize for anyone using AI for investment research is that AI is by definition a consensus machine .... it is up to you to determine the steel man before involving AI.
Multi-agent workflows changed what one operator can produce
The EQDP dossier wasn’t a single ChatGPT conversation. It was Claude Code orchestrating multiple sub-agents in parallel.
That kind of structured, multi-thread research used to require a small team. Now one operator with the right framework can produce institutional-depth output in an afternoon.
The implication for the buy-side is straightforward. The marginal cost of producing the details going into a research dossier is approaching zero, while the marginal cost of writing a good research hypothesis has not.
Practical Takeaway
If you're using AI for investing research, spend 80% of your time on the framework: the dimensions, the lenses, the disqualifiers. That is the work no one else can copy.
Pair that with a multi-agent workflow that can fill out the details, and you have a research process that compresses weeks into hours without giving up rigor.
Personal
Speaking of Singapore, I'm there in May running workshops for our hedge fund clients and speaking at a few sell-side events.
If you're based there and want to come, hit reply and I'll try to get you access.
Talk soon,
-Dave
When you're ready, here's how we can work together: