The AI Can't be Wrong and Neither Can I

Train on human data, make human mistakes

· 5 min read

I have been working on a somewhat simple trading bot. No sense going into too many details except to say the concept is fairly simple - when price goes up enough in the short-term it likely reverts to the mean. Decided to do some historical testing using Bitcoin pricing data to see how the idea fared historically. Claude Opus 5.5 kept telling me the idea lost money over about a six-year period. Claude may be right, but it didn’t sound right to me. So, I asked Claude to try to change some variables. Maybe tweak the target price or change the stop loss. Set the stop loss and target price equidistant from the starting price. Try different Average True Ranges. Just try different things. No dice. Claude was convinced that no matter what, the bot would lose money. OK. The thing eating at me was that two out of three trades the bot had completed were winners and about the same ratio of trades I performed manually also won. Sure, small sample sizes and all of that. But, I wasn’t convinced that Claude’s conclusion that less than one third of trades would be successful based on its historical analysis.

Me being me, I decide to let Codex take a crack at it. I explained that I thought Claude was wrong, but that it is always possible I am incorrect. I passed on information about how Claude had conducted its historical tests and asked Codex to look at how the bot was actually conducting trades, to take a look at the completed trades, and to re-run Claude’s analysis. Same or very similar results. I began to point out the errors in Claude’s methodology. It was using a proxy for price data and not the actual source of price data the bot was using. The difference between the prices of the proxy and the price data source the bot used was enough to cause the analysis to be incorrect. Codex still concluded Claude was basically correct. OK, fine. I asked Codex to look at where the stops and targets were placed. Was there some different combination of stops and targets that would prove more successful. Codex concluded even changing those that over every period the trades would lose. It’s possible, but it still didn’t sound right to me.

Not being one to give up I asked Codex to re-run the analysis from scratch. Disregard whatever Claude had come up with. Create your own test and just do it over again. Codex did that and amazingly concluded Claude was essentially correct. Again I asked Codex to consider where stops were set and target prices. Why not look at setting target price and stop losses equidistant from some percentage of the distance between bands of standard deviation? This back and forth went on for a while. I probably went through about a dozen versions total of Claude and Codex saying that absolutely, under almost any time period during the last six years, the trading bot strategy would definitely lose money. In fact, according to both of them, the trading strategy was so lousy that it lost money before any fees. The average person would probably give up after the first several quite convincing arguments backed by data. But I am not the average person. I am the stubborn person who has questions. Lots and lots of questions.

After bouncing questions and answers back and forth for even longer, Codex finally came back with a potential winner - ATR (Average True Range) - somewhere between two and three. I noted that is a wide range and could Codex narrow it down which it did. 2 - 2.25 and 2.75 to 3 appeared to make money with 2 - 2.25 having an edge. Now, I appreciate abstractions and acronyms, but for Bitcoin I wanted to know what kind of distance for a stop and target distance from entry were we talking about. It ended up being around $2,000 which is almost exactly where I landed on during manually trading.

So after arguing over historical reconstructions of trade for several hours and being told repeatedly that under any scenario imaginable that the bot would most definitely lose money and me arguing that I did not agree Codex found a stop loss and target price distance that almost exactly matched what I came up fiddling around on my own. After a while it feels like gaslighting, except it isn’t. Claude and Codex aren’t human. They have no experience. They have no knowledge, They have no intuition. They are incredible coders and can also create incredibly brilliant code. But they easily get stuck in a loop. They reach a conclusion based on some narrow slice of data - often received from me or some other human - and conclude that their conclusion is correct and no amount of questioning or re-examination of the data can prove anything but the original answer. The trading bot will always lose money regardless of market and not matter how the stops or target prices are set.

None of it makes sense. The conclusions pulled from studying the data and supposedly following the same trading method as the bot led to the opposite conclusion from what I was seeing with my own eyes. And, believe me, I am not going to just believe my lying eyes when trading, but sometimes your eyes and your gut are all you have after analysis and all else fails. Even after all of that, I didn’t change anything about the trading bot at all. It has barely done any trades and just because I think I am right and possibly bulldogged Codex into agreeing with me doesn’t mean I am correct. Just because Claude and Codex found a dozen ways to tell me the bot will definitely lose money doesn’t make them right. In the end, there is only evidence of the bot results over a long period of time. It will have to trade over bull, bear and sideways markets. Eventually, it may prove Claude and Codex right or me right. Or, it may prove all of us to be fools and trade flat. Nobody knows. I think I know. Claude and Codex run the numbers so they obviously know (but they don’t know). Nobody knows. The only way we’ll ever know is to be patient and keep letting the bot run. Fix mistakes in the code. Repeat. That’s it.