Article Contents
REVIEW   Open Access     Cite

Context Scaling: A New Frontier of AI Scaling

More Information
  • Corresponding author: xpqiu@fudan.edu.cn
  • DownLoad: Full size image
    1. CGS Some deployed failures arise because current task state is unavailable, distorted, poorly selected, or unused.

      Context scaling is a research program for measuring returns on context-infrastructure budget Bc.

      Training, decision-time, and context budgets interact and should be reported separately.

      Effective-context value is measured after intervention against a declared task and baseline.

      Utility-cost curves, not window length alone, are the primary measurement target.

  • A recurring failure mode in deployed AI lies at the boundary between the state a task depends on and the model-visible context available when the system acts. In such cases, what looks like weak reasoning can instead reflect failed access: relevant user state, repository state, tool traces, sensor readings, or environment changes are absent, distorted, or poorly selected at decision time. Viewing a model as a conditional Pθ (d|c) embedded in a system with three budgets (training budget Bθ, decision-time budget Bd, and context-infrastructure budget Bc), we read recent progress as a sequence of bottleneck shifts. Pretraining scaling spends Bθ to improve the base model; decision-time scaling spends Bd to search, refine, or verify candidate decisions under a supplied context; context scaling spends Bc to improve what the system can make available before a decision. The three programs interact. We treat context infrastructure as a distinct investment and measurement target, while leaving the form of any general scaling law open. Reporting only the model and per-query inference compute obscures the contribution of context infrastructure. The relevant unit of analysis is the model--infrastructure pair. Stronger models can exploit richer decision-time information states; stronger context infrastructure can acquire, represent, deliver, evaluate, and persist those states; infrastructure-mediated deployment traces can improve both. Effective-context value is a task- and intervention-dependent attribution measured after observing the utility of model-visible context. Availability, fidelity, selection, and utilization failures separate potentially relevant deployment state from information that improves a decision. The paper closes with three problem clusters: grounded evaluation and reporting, interaction-data supply and full-cost infrastructure, and safety and governance for persistent and adaptive systems.
  • 加载中
  • [1] Shapley L. S. (1953). A Value for n-Person Games. Contributions to the Theory of Games 2:307−17. DOI:10.7249/p0295

    View in Article CrossRef Google Scholar

    [2] Lundberg S. M. and Lee S. I. (2017). A unified approach to interpreting model predictions. Advances in neural information processing systems 4765−74.

    View in Article Google Scholar

    [3] OpenAI (2024). Learning to reason with LLMs. (OpenAI public report; https://openai.com/index/learning-to-reason-with-llms/).

    View in Article Google Scholar

    [4] DeepSeek-AI, Guo D., Yang D., et al. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645:633−8. DOI:10.1038/s41586-025-09422-z

    View in Article CrossRef Google Scholar

    [5] Jin B., Zeng H., Yue Z., et al. (2025). Search-R1: Training LLMs to reason and leverage search engines with reinforcement learning. Conference on language modeling.

    View in Article Google Scholar

    [6] Kaplan J., McCandlish S., Henighan T., et al. (2020). Scaling laws for neural language models. arXiv preprint arXiv: 2001.08361.

    View in Article Google Scholar

    [7] Hoffmann J., Borgeaud S., Mensch A., et al. (2022). An empirical analysis of Compute-Optimal large language model training. Advances in neural information processing systems.

    View in Article Google Scholar

    [8] Muennighoff N., Rush A. M., Barak B., et al. (2023). Scaling Data-Constrained language models. Advances in neural information processing systems.

    View in Article Google Scholar

    [9] Sorscher B., Geirhos R., Shekhar S., et al. (2022). Beyond neural scaling laws: Beating power law scaling via data pruning. Advances in neural information processing systems.

    View in Article Google Scholar

    [10] Wei J., Wang X., Schuurmans D., et al. (2022). Chain-of-Thought prompting elicits reasoning in large language models. Advances in neural information processing systems.

    View in Article Google Scholar

    [11] Wang X., Wei J., Schuurmans D., et al. (2023). Self-Consistency improves chain of thought reasoning in language models. International conference on learning representations.

    View in Article Google Scholar

    [12] Yao S., Yu D., Zhao J., et al. (2023). Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems.

    View in Article Google Scholar

    [13] Lightman H., Kosaraju V., Burda Y., et al. (2024). Let’s verify step by step. International conference on learning representations.

    View in Article Google Scholar

    [14] Cobbe K., Kosaraju V., Bavarian M., et al. (2021). Training verifiers to solve math word problems. arXiv preprint arXiv: 2110.14168.

    View in Article Google Scholar

    [15] Snell C., Lee J., Xu K., et al. (2025). Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning. International conference on learning representations.

    View in Article Google Scholar

    [16] Wei Z., Yao W., Liu Y., et al. (2025). WebAgent-R1: Training web agents via End-to-End Multi-Turn reinforcement learning. Conference on empirical methods in natural language processing 7909−28. DOI:10.18653/v1/2025.emnlp-main.401

    View in Article CrossRef Google Scholar

    [17] Zou D. and others (2026). On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents. arXiv preprint arXiv: 2603.12109.

    View in Article Google Scholar

    [18] Montgomery K., Park D., Tu J., et al. (2025). Predicting Task Performance with Context-aware Scaling Laws. arXiv preprint arXiv: 2510.14919.

    View in Article Google Scholar

    [19] Shi J., Ma Q., Liu H., et al. (2025). Intrinsic entropy of context length scaling in LLMs. arXiv preprint arXiv: 2502.01481.

    View in Article Google Scholar

    [20] Fang Y., Zhan J., Ai Q., et al. (2024). Scaling laws for dense retrieval. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval.

    View in Article Google Scholar

    [21] Fang L., Wang Y., Liu Z., et al. (2025). What is Wrong with Perplexity for Long-context Language Modeling? International conference on learning representations.

    View in Article Google Scholar

    [22] Kimi Team, Du A., Gao B., et al. (2025). Kimi k1.5: Scaling reinforcement learning with LLMs. arXiv preprint arXiv: 2501.12599.

    View in Article Google Scholar

    [23] DeepSeek-AI (2026). DeepSeek-V4: Towards highly efficient Million-Token context intelligence. (Hugging Face model card / technical report; https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro).

    View in Article Google Scholar

    [24] GLM-4.5 Team, Zeng A., Lv X., et al. (2025). GLM-4.5: Agentic, reasoning, and coding (ARC) foundation models. arXiv preprint arXiv: 2508.06471.

    View in Article Google Scholar

    [25] Z.ai (2026). GLM-4.7 & GLM-4.6 & GLM-4.5. (GitHub repository; https://github.com/zai-org/GLM-4.5).

    View in Article Google Scholar

    [26] Hsieh C. P., Sun S., Kriman S., et al. (2024). RULER: What’s the real context size of your Long-Context language models? arXiv preprint arXiv: 2404.06654.

    View in Article Google Scholar

    [27] Bai Y., Tu S., Zhang J., et al. (2025). LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks. Annual meeting of the association for computational linguistics 3639−64. DOI:10.18653/v1/2025.acl-long.183

    View in Article CrossRef Google Scholar

    [28] Shao R., He J., Asai A., et al. (2024). Scaling Retrieval-Based language models with a Trillion-Token datastore. Advances in neural information processing systems.

    View in Article Google Scholar

    [29] Yue Z., Zhuang H., Bai A., et al. (2025). Inference scaling for Long-Context retrieval augmented generation. International conference on learning representations.

    View in Article Google Scholar

    [30] Hernandez D., Kaplan J., Henighan T., et al. (2021). Scaling laws for transfer. arXiv preprint arXiv: 2102.01293.

    View in Article Google Scholar

    [31] Sutton R. S. (2019). The bitter lesson. (Essay; http://www.incompleteideas.net/IncIdeas/BitterLesson.html).

    View in Article Google Scholar

    [32] Xing E., Deng M. and Hou J. (2026). Critique of agent model. arXiv preprint arXiv: 2606.23991.

    View in Article Google Scholar

    [33] OpenAI (2025). Introducing deep research. (OpenAI product announcement; https://openai.com/index/introducing-deep-research/).

    View in Article Google Scholar

    [34] Mialon G., Fourrier C., Swift C., et al. (2024). GAIA: A benchmark for general AI assistants. International conference on learning representations.

    View in Article Google Scholar

    [35] Wei J., Sun Z., Papay S., et al. (2025). BrowseComp: A simple yet challenging benchmark for browsing agents. (OpenAI technical report; https://openai.com/index/browsecomp/).

    View in Article Google Scholar

    [36] Silver D. and Sutton R. S. (2025). Welcome to the era of experience. (Preprint of a chapter in Designing an Intelligence; https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf).

    View in Article Google Scholar

    [37] Karpathy A. (2025). On context engineering over prompt engineering. (X (Twitter) post; https://x.com/karpathy/status/1937902205765607626).

    View in Article Google Scholar

    [38] LangChain (2025). Context engineering for agents. (LangChain blog post; https://blog.langchain.com/context-engineering-for-agents/).

    View in Article Google Scholar

    [39] He C., Zhou X., Wang D., et al. (2026). Harness engineering for language agents: The harness layer as control, agency, and runtime. Preprints. DOI: 10.20944/preprints202603.1756.v2

    View in Article Google Scholar

    [40] Meng Q., Wang Y., Chen L., et al. (2026). Agent harness for large language model agents: A survey. Preprints. DOI: 10.20944/preprints202604.0428.v3

    View in Article Google Scholar

    [41] Lewis P., Perez E., Piktus A., et al. (2020). Retrieval-Augmented generation for Knowledge-Intensive NLP tasks. Advances in neural information processing systems.

    View in Article Google Scholar

    [42] Schick T., Dwivedi-Yu J., Dessì R., et al. (2023). Toolformer: Language models can teach themselves to use tools. Advances in neural information processing systems.

    View in Article Google Scholar

    [43] Yao S., Zhao J., Yu D., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International conference on learning representations.

    View in Article Google Scholar

    [44] Yan S., Yang X., Huang Z., et al. (2026). Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning. Annual meeting of the association for computational linguistics 12805−25.

    View in Article Google Scholar

    [45] Zhou C., Chai H., Chen W., et al. (2026). Externalization in LLM agents: A unified review of memory, skills, protocols and harness engineering. arXiv preprint arXiv: 2604.08224.

    View in Article Google Scholar

    [46] Sutton R. S. and Barto A. G. (2018). Reinforcement learning: An introduction. (MIT Press).

    View in Article Google Scholar

    [47] Liu X., Li R., Huang M., et al. (2025). Thus spake Long-Context large language model. arXiv preprint arXiv: 2502.17129.

    View in Article Google Scholar

    [48] Kaelbling L. P., Littman M. L. and Cassandra A. R. (1998). Planning and acting in partially observable stochastic domains. Artif Intell 101:99−134. DOI:10.1016/S0004-3702(98)00023-X

    View in Article CrossRef Google Scholar

    [49] Bajcsy R. (1988). Active perception. Proc. IEEE 76:996−1005. DOI:10.1109/5.5968

    View in Article CrossRef Google Scholar

    [50] Marchionini G. (2006). Exploratory search: From finding to understanding. Commun ACM 49:41−6. DOI:10.1145/1121949.1121979

    View in Article CrossRef Google Scholar

    [51] Hollan J., Hutchins E. and Kirsh D. (2000). Distributed cognition: Toward a new foundation for Human-Computer interaction research. ACM Trans. Comput.-Hum. Interact. 7:174−96. DOI:10.1145/353485.353487

    View in Article CrossRef Google Scholar

    [52] Todorov E., Erez T. and Tassa Y. (2012). MuJoCo: A physics engine for model-based control. 2012 IEEE/RSJ international conference on intelligent robots and systems (IEEE): 5026–33.

    View in Article Google Scholar

    [53] Ha D. and Schmidhuber J. (2018). World models. arXiv preprint arXiv: 1803.10122.

    View in Article Google Scholar

    [54] Hafner D., Pasukonis J., Ba J., et al. (2025). Mastering diverse control tasks through world models. Nature 640:647−53. DOI:10.1038/s41586-025-08744-2

    View in Article CrossRef Google Scholar

    [55] Défossez A., Mazaré L., Orsini M., et al. (2024). Moshi: a speech-text foundation model for real-time dialogue. arXiv preprint arXiv: 2410.00037.

    View in Article Google Scholar

    [56] Liu H., Li C., Wu Q., et al. (2023). Visual instruction tuning. Advances in neural information processing systems.

    View in Article Google Scholar

    [57] Xu Y., Li M., Cui L., et al. (2020). LayoutLM: Pre-training of Text and Layout for Document Image Understanding. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining.

    View in Article Google Scholar

    [58] Kim M. J., Pertsch K., Karamcheti S., et al. (2024). OpenVLA: An Open-Source Vision-Language-Action model. Conference on robot learning 2679−713.

    View in Article Google Scholar

    [59] Cursor (2026). Training composer for longer horizons. (https://cursor.com/blog/self-summarization).

    View in Article Google Scholar

    [60] Chroma (2026). Chroma context-1: Training a Self-Editing search agent. (https://www.trychroma.com/research/context-1).

    View in Article Google Scholar

    [61] Wu Q., Bansal G., Zhang J., et al. (2024). AutoGen: Enabling Next-Gen LLM applications via Multi-Agent conversation. Proceedings of the first conference on language modeling.

    View in Article Google Scholar

    [62] Sun W., Lu M., Ling Z., et al. (2025). Scaling Long-Horizon LLM agent via Context-Folding. arXiv preprint arXiv: 2510.11967.

    View in Article Google Scholar

    [63] Liu X., Yan H., An C., et al. (2024). Scaling Laws of RoPE-based Extrapolation. International conference on learning representations.

    View in Article Google Scholar

    [64] Shi L., Zhang H., Yao Y., et al. (2024). Keep the cost down: A review on methods to optimize LLM’s KV-Cache consumption. arXiv preprint arXiv: 2407.18003.

    View in Article Google Scholar

    [65] Xiao G., Tian Y., Chen B., et al. (2024). Efficient streaming language models with attention sinks. International conference on learning representations.

    View in Article Google Scholar

    [66] Liu N. F., Lin K., Hewitt J., et al. (2024). Lost in the middle: How language models use long contexts. Trans Assoc Comput Linguist.

    View in Article Google Scholar

    [67] Lin J., Liu S., Pan C., et al. (2026). Agentic harness engineering: Observability-Driven automatic evolution of Coding-Agent harnesses. arXiv preprint arXiv: 2604.25850.

    View in Article Google Scholar

    [68] Amodei D., Olah C., Steinhardt J., et al. (2016). Concrete problems in AI safety. arXiv preprint arXiv: 1606.06565.

    View in Article Google Scholar

    [69] Pan A., Bhatia K. and Steinhardt J. (2022). The effects of reward misspecification: Mapping and mitigating misaligned models. International conference on learning representations.

    View in Article Google Scholar

    [70] Packer C., Wooders S., Lin K., et al. (2023). MemGPT: Towards LLMs as operating systems. arXiv preprint arXiv: 2310.08560.

    View in Article Google Scholar

    [71] Zhang Z., Bo X., Ma C., et al. (2025). A Survey on the Memory Mechanism of Large Language Model based Agents. ACM Trans Inf Syst 43:155:1−155:47. DOI:10.1145/3748302

    View in Article CrossRef Google Scholar

    [72] Yu Y., Yao L., Xie Y., et al. (2026). Agentic memory: Learning unified Long-Term and Short-Term memory management for large language model agents. Annual meeting of the association for computational linguistics 21457−83.

    View in Article Google Scholar

    [73] Shavit Y., Agarwal S., Brundage M., et al. (2023). Practices for governing agentic AI systems. (OpenAI research paper; https://cdn.openai.com/papers/practices-for-governing-agentic-ai-systems.pdf).

    View in Article Google Scholar

    [74] Casper S., Ezell C., Siegmann C., et al. (2024). Black-Box Access is Insufficient for Rigorous AI Audits. Proceedings of the 2024 ACM conference on fairness, accountability, and transparency.

    View in Article Google Scholar

    [75] Ning X., Tieu K., Fu D., et al. (2026). Code as agent harness. arXiv preprint arXiv: 2605.18747.

    View in Article Google Scholar

    [76] Pan L., Zou L., Guo S., et al. (2026). Natural-Language agent harnesses. arXiv preprint arXiv: 2603.25723.

    View in Article Google Scholar

    [77] OpenClaw (2026). Session management: Compaction. (https://open-claw.bot/docs/cli/reference/session-management-compaction/).

    View in Article Google Scholar

    [78] Jimenez C. E., Yang J., Wettig A., et al. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? International conference on learning representations.

    View in Article Google Scholar

    [79] Zhou X., Zhu H., Mathur L., et al. (2024). SOTOPIA: Interactive evaluation for social intelligence in language agents. International conference on learning representations.

    View in Article Google Scholar

    [80] Fu Y., Peng H., Khot T., et al. (2023). Improving language model negotiation with Self-Play and In-Context learning from AI feedback. arXiv preprint arXiv: 2305.10142.

    View in Article Google Scholar

    [81] Arora S., Lu Z., Chiu C. C., et al. (2025). Talking turns: Benchmarking audio foundation models on Turn-Taking dynamics. International conference on learning representations.

    View in Article Google Scholar

    [82] Google (2026). New AI tools for the future of science. (https://blog.google/innovation-and-ai/technology/research/gemini-for-science-io-2026/).

    View in Article Google Scholar

    [83] Gottweis J., Weng W. H., Daryin A., et al. (2026). Accelerating scientific discovery with Co-Scientist. Nature. DOI: 10.1038/s41586-026-10644-y

    View in Article Google Scholar

    [84] Aygün E., Belyaeva A., Comanici G., et al. (2026). An AI system to help scientists write expert-level empirical software. Nature. DOI: 10.1038/s41586-026-10658-6

    View in Article Google Scholar

    [85] Li R., Zhang X., Yu H., et al. (2026). MemPO: Self-Memory policy optimization for Long-Horizon agents. arXiv preprint arXiv: 2603.00680.

    View in Article Google Scholar

    [86] Zhou S., Xu F. F., Zhu H., et al. (2024). WebArena: A realistic web environment for building autonomous agents. International conference on learning representations.

    View in Article Google Scholar

    [87] Liu X., Yu H., Zhang H., et al. (2024). AgentBench: Evaluating LLMs as agents. International conference on learning representations.

    View in Article Google Scholar

    [88] Maharana A., Lee D. H., Tulyakov S., et al. (2024). Evaluating very Long-Term conversational memory of LLM agents. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics.

    View in Article Google Scholar

    [89] Zhong W., Guo L., Gao Q., et al. (2024). MemoryBank: Enhancing large language models with Long-Term memory. AAAI conference on artificial intelligence 19724−31. DOI:10.1609/aaai.v38i17.29946

    View in Article CrossRef Google Scholar

    [90] Zhao W., Queralta J. P. and Westerlund T. (2020). Sim-to-Real transfer in deep reinforcement learning for robotics: a survey. 2020 IEEE symposium series on computational intelligence (SSCI).

    View in Article Google Scholar

    [91] Li X., Hsu K., Gu J., et al. (2024). Evaluating Real-World robot manipulation policies in simulation. Conference on robot learning 3705−28.

    View in Article Google Scholar

    [92] Abadi M., Chu A., Goodfellow I., et al. (2016). Deep learning with differential privacy. Proceedings of the 2016 ACM SIGSAC conference on computer and communications security.

    View in Article Google Scholar

    [93] Shumailov I., Shumaylov Z., Zhao Y., et al. (2024). AI models collapse when trained on recursively generated data. Nature 631:755−9. DOI:10.1038/s41586-024-07566-y

    View in Article CrossRef Google Scholar

    [94] Sardana N., Portes J., Doubov S., et al. (2024). Beyond Chinchilla-Optimal: Accounting for inference in language model scaling laws. Proceedings of the 41st International Conference on Machine Learning.

    View in Article Google Scholar

    [95] Greshake K., Abdelnabi S., Mishra S., et al. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec) 79−90.

    View in Article Google Scholar

    [96] Perez E., Huang S., Song F., et al. (2022). Red teaming language models with language models. Proceedings of the 2022 conference on empirical methods in natural language processing.

    View in Article Google Scholar

    [97] Ilyas A., Park S. M., Engstrom L., et al. (2022). Datamodels: Understanding predictions with data and data with predictions. International conference on machine learning 9525−87.

    View in Article Google Scholar

    [98] Coalition for Content Provenance and Authenticity (C2PA) (2024). C2PA technical specifications. (Open technical standard; https://spec.c2pa.org/specifications/specifications/2.2/index.html).

    View in Article Google Scholar

  • Cite this article:

    Qiu X. (2026). Context Scaling: A New Frontier of AI Scaling. AI Plus 1:100010. https://doi.org/10.59717/ipj.aiplus.2026.100010
    Qiu X. (2026). Context Scaling: A New Frontier of AI Scaling. AI Plus 1:100010. https://doi.org/10.59717/ipj.aiplus.2026.100010

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(3)     Tables(5)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(337) PDF downloads(71)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint