The Frontier
Declarative Attention cuts long-context KV reads by up to 52%
A paper from KAIST AI and Google DeepMind argues the model already knows which part of its context matters — and a community fork just put that claim into llama.cpp. An off-the-shelf model can be prompted to declare, inside its own chain of thought, which slice of its