2026-03-10: Post-AGI safety
I am currently working on state-of-the-art multimodal reasoning models to better understand their limitations. They are becoming very smart at a rapid pace.
For the first time in human history (at least in recorded history), we may soon have agents with both better thinking abilities and stronger physical capabilities than our own that may find ways to survive without us.
Some people argue that this will bring great prosperity, allowing humans to work less and enjoy life more. However, I believe that, at the very least, we will still need to work on controlling these agents—ensuring that they benefit us rather than harm us—and this task will become increasingly difficult.
If history is any guide, the natural outcome of such an imbalance is that weaker agents may gradually become extinct, much as many animal species did when humans came to dominate them.
On the other hand, humans have been shown to become smarter when competing with strong opponents. Whether this dynamic will reduce the risks above remains uncertain; based on the current trajectory, I am not very optimistic.
We should pay much more attention to this before it is too late.