Where robotics and AI technologies came from, which approaches now compete or converge, and what problems remain unresolved.
This is an editorially selected reading order. Begin at step 1 or go directly to the topic you need.
From filter-based methods to graph optimization, and now to learning-based, AI-native approaches. Tracing the lineage of SLAM (simultaneous self-localization and mapping) and where current research stands.
From the LOAM lineage to direct methods like FAST-LIO, and on to D-LIO, published in 2026. Surveying the evolution of LiDAR-and-IMU-based self-localization and mapping, set against newbot's own implementation.
From classical feature-point methods to deep-learning-based end-to-end estimation, and on to a reinvention of map representation via 3D Gaussian Splatting/NeRF. Surveying where Visual-SLAM — self-localization and mapping from cameras alone — stands today.
The YOLO family moving toward NMS-free designs, and transformer-based detectors like RF-DETR, which broke 60 mAP on COCO. Surveying the state of object detection in 2026, all the way to open-vocabulary detection driven by nothing but a text prompt.
From transformer-based SegFormer and Mask2Former to the Segment Anything family, which can carve out any object from a prompt. Surveying where segmentation stands as it moves away from fixed classes.
Ollama's monthly download count grew 520x in three years. Surveying how the rapid rise of open-weight models and maturing quantization techniques are turning large language models into something that runs on a home PC.
A 'World Model' learns the dynamics of an environment and simulates the future inside its own head. From Ha & Schmidhuber's VAE+RNN, through DreamerV3, Genie 3, and Sora, to the JEPA architecture championed by Yann LeCun — this piece maps the split between generative and non-generative approaches, and the shared wall of physical consistency both sides run into.
Vision-Language-Action folds camera input and natural-language instructions into a single model that outputs robot actions directly. From RT-2 to OpenVLA to Physical Intelligence's π0 lineage, surveying the spread of a design philosophy that refuses to separate perception from control.
Tesla Optimus and Figure 03 have moved onto factory floors, and Boston Dynamics' Atlas has switched from hydraulic to fully electric. Unitree is driving a price war, while 1X's NEO is starting to teach itself in the home using a video world model. A look at the competitive landscape and the open problems that remain.