They trained models on only task specific data, not on a general dataset and certainly not on the enormous datasets frontier models are trained on.
"Our training sets consist of 2.9M sequences (120M tokens) for shortest paths; 31M sequences (1.7B tokens) for noisy shortest paths; and 91M sequences (4.7B tokens) for random walks. We train two types of transformers [38] from scratch using next-token prediction for each dataset: an 89.3M parameter model consisting of 12 layers, 768 hidden dimensions, and 12 heads; and a 1.5B parameter model consisting of 48 layers, 1600 hidden dimensions, and 25 heads."
Or, ignore the hype, look at what we know about how these models work and about the structures their weights represent, and base your answers on that today.
https://neurosciencenews.com/llm-ai-logic-27987/