IBM announced the release of its Granite 4.2 language models on August 25, 2026, marking a shift towards models designed for agentic workloads and predictable enterprise deployment. The new family includes models with 3 billion, 8 billion, and 30 billion parameters, all released under the Apache 2.0 license. This licensing allows organizations to download, fine-tune, and deploy the models in commercial production without licensing restrictions. A core feature of Granite 4.2 is its native reasoning ability, which enables models to perform step-by-step "chain-of-thought" processing before generating a final answer. This capability is exposed through a "thinking" switch in the chat template, alongside "low-effort" and "non-thinking" modes, providing users with control over the depth of reasoning and latency per query. This design aims to improve performance on complex tasks such as mathematics, coding, and multi-step logic. The models are decoder-only dense transformers, pre-trained on approximately 15 trillion tokens. The training regimen for Granite 4.2 involved a multi-stage reinforcement learning process. For the 8B and 30B parameter models, this included an agentic reinforcement learning block where the models learned to edit code, operate a terminal, and conduct web searches within sandboxed environments. This agentic training is intended to equip the models for complex, multi-step enterprise tasks. IBM emphasizes the deployability of Granite 4.2 across various environments, including cloud, on-premises, and edge devices. The dense architecture supports broad compatibility, and the range of model sizes offers flexibility for different computational requirements. The 3B model is designed for solo developers and startups, while the 8B model suits mid-market teams, and the 30B model is intended for enterprises with substantial GPU capacity. Granite 4.2 also incorporates reasoning-augmented tool calling. The models are designed to evaluate which tools to invoke and why, before executing a call in an OpenAI function-calling format. This allows for integration into existing agentic harnesses without requiring additional adapters. The models are tested across 12 languages, including English, German, Japanese, Arabic, Korean, and Chinese. In addition to the language models, IBM also released two 470-million-parameter Granite Speech 5.0 Turbo CTC models. These speech models are designed for efficient streaming audio transcription and are optimized for deployment on edge devices due to their compact size and lack of an LLM backbone. IBM reports a throughput of approximately 12,600 RTFx on a single H200 GPU for these speech models. The development of Granite 4.2 included training on 1 trillion tokens of synthetic code generated via IBM's CodeAlchemy pipeline. An intermediate training step, referred to as "mid-training," was also employed to enhance reasoning capabilities. The models feature a speculative decoding layer to improve text output speed and reduce operating costs for enterprises. IBM is also collaborating with Hirundo to use machine unlearning technology to reduce undesirable model outputs without full retraining. IBM's strategy for its open models involves a division of labor. While earlier Granite generations focused on instruction-following, Granite 4.2 is positioned as a reasoning-first agent for enterprise applications. The company has made the full training recipe, data-mixture proportions, and per-stage hyperparameters publicly available alongside the model weights.
IBM Releases Granite 4.2 Models with Enhanced Reasoning and Agentic Capabilities
IBM has launched its Granite 4.2 family of large language models, featuring explicit reasoning capabilities and agentic reinforcement learning for enterprise deployments. These open-source models are available in 3B, 8B, and 30B parameter sizes under an Apache 2.0 license, allowing for flexible on-premise and cloud integration.
AI-assisted archive. Original-source fact review is pending. How evidence is checked
Read source-checked investigations