For anyone interested in digging deeper on LLM navigation capabilities, there is the work I mentioned in the post[0], and also a paper that tried to test this more formally in 2024[1]. I didn't learn of these until after I did my experiment, it would have been interesting to retest the same routes.
Maybe in a year, when everything has changed once again.
Maybe in a year, when everything has changed once again.
[0] https://github.com/Hansenq/nav-evals-public [1] https://arxiv.org/abs/2411.17912