Бизнес и продуктивность

Масштабирование применения ИИ: как меняется экономика и архитектура вычислений

К 2030 году на применение ИИ будет приходиться около 60% вычислительного спроса. Индустрия переходит от создания моделей к их обслуживанию, что требует новых аппаратных архитектур.

Обновлено:
7 мин чтения
Масштабирование применения ИИ: как меняется экономика и архитектура вычислений

Суть

Развертывание инфраструктуры искусственного интеллекта прошло важный рубеж. Долгие годы большая часть вычислительных ресурсов направлялась на обучение моделей, что требовало создания кластеров из тысяч графических процессоров (GPU). Однако, по прогнозам McKinsey, к 2030 году структура рабочих нагрузок изменится: на применение моделей (inference) будет приходиться около 60% спроса на ИИ, тогда как на обучение — около 40%. Индустрия переходит от создания интеллекта к его масштабному обслуживанию, что представляет собой совершенно иную инженерную задачу с другой структурой затрат.

Контекст

Руководители компаний больше не задаются вопросом, создает ли ИИ ценность. Главный вопрос сегодня — делает ли он это при приемлемой стоимости обслуживания. Внедрение ИИ требует немедленных затрат на инфраструктуру, тогда как перестройка процессов и получение измеримого роста производительности занимают время. Если стоимость обслуживания не снизится значительно, предприятия просто не смогут применять ИИ в больших масштабах.

Чтобы понять эту новую среду, McKinsey совместно с Global Semiconductor Alliance (GSA) провела интервью с архитекторами ИИ из ведущих полупроводниковых компаний и облачных провайдеров (hyperscalers).

Image description: A stacked bar chart shows demand by workload type under the continued momentum scenario. The total demand in 2026 was 120 gigawatts, of which AI training accounted for 29% of the total, AI inference 29%, and non-AI 42%. In 2030, the total demand is expected to be 278 gigawatts, with AI training making up 28% of the total, AI inference 42%, and non-AI 30%. AI inference is expected to make up 60% of all AI demand in 2030, with a per annum growth of 36%. Total AI growth is expected to be 30% per annum. End image description.

Image description: A stacked bar chart shows demand by workload type under the continued momentum scenario. The total demand in 2026 was 120 gigawatts, of which AI training accounted for 29% of the total, AI inference 29%, and non-AI 42%. In 2030, the total demand is expected to be 278 gigawatts, with AI training making up 28% of the total, AI inference 42%, and non-AI 30%. AI inference is expected to make up 60% of all AI demand in 2030, with a per annum growth of 36%. Total AI growth is expected to be 30% per annum. End image description.

Детали

Применение ИИ фундаментально неоднородно. Требования к задержке (latency), пропускной способности и стоимости различаются в зависимости от задачи, будь то голосовой помощник реального времени или анализ 300-страничного контракта.

Более того, требования различаются даже внутри одного запроса. Фаза предварительного заполнения (prefill), когда модель обрабатывает подсказку, требует высоких вычислительных мощностей. Фаза декодирования (decode), когда генерируются токены, зависит в первую очередь от пропускной способности памяти. За последние два десятилетия пиковая вычислительная мощность росла примерно в два раза быстрее, чем пропускная способность памяти.

Большая часть современной инфраструктуры (включая конфигурации GPU и памяти с высокой пропускной способностью HBM) была разработана для обучения, которое вознаграждает чистую вычислительную мощность. Использование этого же оборудования для декодирования создает дисбаланс, так как генерация каждого токена требует выгрузки весов модели из памяти при сравнительно небольшом объеме арифметических операций.

Анализ

Image description: A table shows workload requirements and estimated share of 2030 AI inference token demand. Under AI training on the table, agentic, code, chat and search, writing, and machine to machine fall under the large language model (LLM) text and code category (82%); media generation and physical AI fall under the non-LLM category (13%); and voice makes up the remainder. Model types across all AI training categories include LLMs, multimodal models, diffusion models, world models, vision models, speech models, recommendation models, and reinforcement models.  Model scale is 10 to 100 billion parameters, up to 128,000-token context window, for agentic, code, and chat and search and is under 10 billion for writing, machine to machine, media generation, physical AI, and voice. Compute is low-latency throughput for agentic and chat and search; high throughput for code, writing, machine to machine, and media generation, and high performance or wattage for physical AI and voice. Memory or networking is high key-value-cache capacity for agentic; high capacity for code; low latency for chat and search, physical AI, and voice; high bandwidth for writing and machine to machine; and moderate memory capacity for media generation.  Serving pattern is as follows: for agentic, iterative agent loop and token by token with cache reuse; for code, long-form token by token with cache reuse; for chat and search, retrieve, then interactive token by token; for writing, token by token with cache reuse; for machine to machine, high-volume request and response; for media generation, long-form token by token with cache reuse; for physical AI, queued jobs, one at a time, 3–50 steps; and for voice, real-time bidirectional streaming. The workload requirement is heavier across model scale, compute, and memory for agentic code, and chat and search. It is moderate for model scale and heavier for compute and memory for writing and machine to machine. It is moderate across model scale and memory for media generation. Last, it is moderate across model scale and compute and heavier in memory for physical AI and voice. Note: Figures may not sum to 100%, because of rounding. Source: AMD; DeepSeek; Google; JEDEC; Lawrence Berkeley National Laboratory; NVIDIA; OpenAI; UALink Consortium; Uptime Institute; McKinsey analysis End image description.

Image description: A table shows workload requirements and estimated share of 2030 AI inference token demand. Under AI training on the table, agentic, code, chat and search, writing, and machine to machine fall under the large language model (LLM) text and code category (82%); media generation and physical AI fall under the non-LLM category (13%); and voice makes up the remainder. Model types across all AI training categories include LLMs, multimodal models, diffusion models, world models, vision models, speech models, recommendation models, and reinforcement models. Model scale is 10 to 100 billion parameters, up to 128,000-token context window, for agentic, code, and chat and search and is under 10 billion for writing, machine to machine, media generation, physical AI, and voice. Compute is low-latency throughput for agentic and chat and search; high throughput for code, writing, machine to machine, and media generation, and high performance or wattage for physical AI and voice. Memory or networking is high key-value-cache capacity for agentic; high capacity for code; low latency for chat and search, physical AI, and voice; high bandwidth for writing and machine to machine; and moderate memory capacity for media generation. Serving pattern is as follows: for agentic, iterative agent loop and token by token with cache reuse; for code, long-form token by token with cache reuse; for chat and search, retrieve, then interactive token by token; for writing, token by token with cache reuse; for machine to machine, high-volume request and response; for media generation, long-form token by token with cache reuse; for physical AI, queued jobs, one at a time, 3–50 steps; and for voice, real-time bidirectional streaming. The workload requirement is heavier across model scale, compute, and memory for agentic code, and chat and search. It is moderate for model scale and heavier for compute and memory for writing and machine to machine. It is moderate across model scale and memory for media generation. Last, it is moderate across model scale and compute and heavier in memory for physical AI and voice. Note: Figures may not sum to 100%, because of rounding. Source: AMD; DeepSeek; Google; JEDEC; Lawrence Berkeley National Laboratory; NVIDIA; OpenAI; UALink Consortium; Uptime Institute; McKinsey analysis End image description.

Эти изменения означают, что предприятия будут оценивать производительность ИИ не по универсальным бенчмаркам, а по стоимости за полезный результат — например, за решенную заявку в службу поддержки или за строку принятого кода.

Изменение метрик может перераспределить ценность на рынке. Сегодня основная прибыль формируется за счет дефицита: передовые ускорители и память HBM продаются с премией. В будущем долгосрочная ценность может перейти к тем уровням цепочки, которые измеримо улучшают экономику всей системы. Это могут быть поставщики памяти, устраняющие узкие места в пропускной способности, разработчики сетевых решений (interconnect) или компании, занимающиеся охлаждением и энергообеспечением.

Перспектива

Разнообразие рабочих нагрузок создает спрос на специализированные чипы. Индустрия уже начинает проектировать решения с учетом разделения задач.

Облачные провайдеры и лаборатории ИИ создают собственные ускорители для самых объемных задач. Компании, делающие ставку на архитектуры на базе SRAM (Groq, Cerebras, SambaNova, d-Matrix), нацеливаются на задачи, ограниченные пропускной способностью памяти. Появляются и комбинированные решения: NVIDIA заключила лицензионное соглашение с Groq, AWS комбинирует чипы Trainium для фазы prefill с решениями Cerebras для decode, AMD сотрудничает с Cerebras, а Intel и SambaNova представили архитектуру, направляющую prefill на GPU, а decode — на реконфигурируемые блоки потоков данных (RDU).

Image description: A table shows the technical readiness and commercial clarity of inference compute location. For the cloud location (for example, frontier model serving) and for the enterprise or co-located location (for example, bank scoring fraud on-premises), they are technically ready and their commercial clarity is settled. For the edge location (for example, visual inspection on the line), it’s technically ready but clarity is unresolved. Last, for the device location (for example, on-phone message summarization), its technical readiness is emerging or use-case dependent and its clarity is partial. End image description.

Image description: A table shows the technical readiness and commercial clarity of inference compute location. For the cloud location (for example, frontier model serving) and for the enterprise or co-located location (for example, bank scoring fraud on-premises), they are technically ready and their commercial clarity is settled. For the edge location (for example, visual inspection on the line), it’s technically ready but clarity is unresolved. Last, for the device location (for example, on-phone message summarization), its technical readiness is emerging or use-case dependent and its clarity is partial. End image description.

A triangle diagram, in which compute, software, and memory and networking make up the three sides of the triangle, shows inference optimization techniques across the three sides. Under compute, techniques are as follows: Custom silicon: Design specialized chips tuned to inference workloads. Distributed computing: Orchestrate model execution (prefill versus decode) across differentiated infrastructure. Advanced nodes (less than 3 nanometers): Achieve energy efficiency and density. Advanced packaging (2.5D or 3D and Chip-on-Wafer-on-Substrate.): Achieve higher bandwidth and chiplet scaling Under software, techniques are as follows: Model optimization techniques (for example, quantization, mixture of experts, distillation, KV-cache optimization) to cut compute and memory cost. Software scheduling techniques (for example, continuous batching, topology-aware scheduling) Last, under memory and network, techniques are as follows: Scale-up interconnect: Use high-speed interconnects to minimize intra-node communication latency, Optical interconnect evolution: Deploy optics-integrated switches (copackaged optics) for higher bandwidth density and lower power per bit, Compute Express Link and dynamic random-access memory tiering: Pool and offload key-value (KV) data to cheaper, shared memory tiers, Low-power double data rate: Integrate low-power memory for energy-efficient inference chips End image description.

A triangle diagram, in which compute, software, and memory and networking make up the three sides of the triangle, shows inference optimization techniques across the three sides. Under compute, techniques are as follows: Custom silicon: Design specialized chips tuned to inference workloads. Distributed computing: Orchestrate model execution (prefill versus decode) across differentiated infrastructure. Advanced nodes (less than 3 nanometers): Achieve energy efficiency and density. Advanced packaging (2.5D or 3D and Chip-on-Wafer-on-Substrate.): Achieve higher bandwidth and chiplet scaling Under software, techniques are as follows: Model optimization techniques (for example, quantization, mixture of experts, distillation, KV-cache optimization) to cut compute and memory cost. Software scheduling techniques (for example, continuous batching, topology-aware scheduling) Last, under memory and network, techniques are as follows: Scale-up interconnect: Use high-speed interconnects to minimize intra-node communication latency, Optical interconnect evolution: Deploy optics-integrated switches (copackaged optics) for higher bandwidth density and lower power per bit, Compute Express Link and dynamic random-access memory tiering: Pool and offload key-value (KV) data to cheaper, shared memory tiers, Low-power double data rate: Integrate low-power memory for energy-efficient inference chips End image description.

This image is a stylized, AI-generated illustration representing technological concepts like artificial intelligence, data processing, or quantum computing. It depicts a central microprocessor or chip with bright blue light emitting upwards, symbolizing high-speed data transmission or intense computational power.

This image is a stylized, AI-generated illustration representing technological concepts like artificial intelligence, data processing, or quantum computing. It depicts a central microprocessor or chip with bright blue light emitting upwards, symbolizing high-speed data transmission or intense computational power.

Colorful iridescent semiconductor wafer with intricate microchip patterns displayed against a dark background.

Colorful iridescent semiconductor wafer with intricate microchip patterns displayed against a dark background.

TL;DR

Главное

Индустрия ИИ смещает фокус с обучения моделей на их применение, что требует перехода от универсальных процессоров к специализированным гибридным решениям из-за высоких затрат на обслуживание.

Ключевые факты

  • /К 2030 году применение моделей составит около 60% спроса на ИИ, а обучение — около 40%.
  • /За последние два десятилетия пиковая вычислительная мощность росла примерно в два раза быстрее, чем пропускная способность памяти.
  • /Компании начинают оценивать эффективность ИИ по стоимости за конкретный результат, а не по стандартным бенчмаркам оборудования.

Инсайт

Долгосрочная ценность на рынке ИИ-оборудования может сместиться от дефицитных вычислительных ускорителей к компонентам, улучшающим общую экономику инфраструктуры — памяти, сетям и системам охлаждения.

Источник:Mckinsey

Читайте также

Гайды по теме