Best Small Language Models (SLMs) - Page 3

How Many Small Language Models (SLMs) Products Does G2 Track?

Total Products under this Category: 40

Category Stats (Sep 2026)

  • Average Rating: 4.24/5 (↓0.02 vs Aug 2026) The average rating of products in this category, based on all submitted ratings
  • Top Trending Product: granite 3.1 MoE 3b (+17.86%) - Among all products in this category, granite 3.1 MoE 3b recorded the largest rating increase compared to last month

Last updated: September 06, 2026

How Does G2 Rank Small Language Models (SLMs) Products?

Why You Can Trust G2's Software Rankings:

  • 30 Analysts and Data Experts
  • 200+ Authentic Reviews
  • 40+ Products
  • Unbiased Rankings

G2's software rankings are built on verified user reviews, rigorous moderation, and a consistent research methodology maintained by a team of analysts and data experts. Each product is measured using the same transparent criteria, with no paid placement or vendor influence. While reviews reflect real user experiences, which can be subjective, they offer valuable insight into how software performs in the hands of professionals. Together, these inputs power the G2 Score, a standardized way to compare tools within every category.

G2 Grid® for Small Language Models (SLMs)

G2 Grid® for  Small Language Models (SLMs)  plotting products by satisfaction and market presence

Highlighted products: Gemma 3 1B, Gemma 3 4B, Gemma 3n 4b, Gemma 3n 2b, Mistral 7B, Gemma 3 270m, Magistral Small, and StableLM.

Underlying data: [Grid® JSON](https://www.g2.com/categories/small-language-models-slms/grids.json?focus%5B%5D=gemma-3-1b&focus%5B%5D=gemma-3-4b&focus%5B%5D=gemma-3n-4b&focus%5B%5D=gemma-3n-2b&focus%5B%5D=mistral-7b&focus%5B%5D=gemma-3-270m&focus%5B%5D=magistral-small&focus%5B%5D=stablelm)

NVIDIA Nemotron Nano 9b

NVIDIA Nemotron-Nano-9B-v2 is a compact, open-source language model designed to deliver high-performance reasoning and agentic capabilities. Utilizing a hybrid Mamba-Transformer architecture, it efficiently processes long-context sequences up to 128,000 tokens, making it suitable for complex tasks requiring extensive context understanding. The model supports multiple languages, including English, German, French, Italian, Spanish, and Japanese, and excels in instruction following and code generation tasks. Key Features and Functionality: - Hybrid Architecture: Combines Mamba-2 state-space layers with Transformer attention layers, enhancing throughput and accuracy in reasoning tasks. - Efficient Long-Context Processing: Capable of handling sequences up to 128,000 tokens on a single NVIDIA A10G GPU, facilitating scalable long-context reasoning. - Multilingual Support: Trained on data spanning 15 languages and 43 programming languages, enabling broad multilingual and coding fluency. - Toggleable Reasoning Feature: Allows users to control the model's reasoning process using simple commands like "/think" or "/no_think," balancing accuracy and response speed. - Reasoning Budget Control: Introduces a "thinking budget" mechanism, enabling developers to set the number of tokens used during the reasoning process, optimizing for latency or cost. Primary Value and User Solutions: NVIDIA Nemotron-Nano-9B-v2 addresses the need for efficient, high-performance language models capable of handling extensive context and complex reasoning tasks. Its hybrid architecture and advanced features provide developers and researchers with a versatile tool for building AI applications that require deep understanding and rapid processing of large-scale textual data. The model's open-source nature and permissive licensing facilitate widespread adoption and customization, empowering users to deploy sophisticated AI solutions across various domains.

Who Is the Company Behind NVIDIA Nemotron Nano 9b?

  • Seller: NVIDIA
  • Year Founded: 1993
  • HQ Location: Santa Clara, CA
  • Twitter: @nvidia
    2,582,827 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    51,762 employees on LinkedIn®
  • Ownership: NVDA

Phi 3.5 mini

Phi-3.5-mini is a lightweight, state-of-the-art language model developed by Microsoft, designed to deliver high-quality reasoning capabilities within a compact architecture. Building upon the datasets used for Phi-3, it focuses on very high-quality, reasoning-dense data, including synthetic data and filtered publicly available websites. The model supports a 128K token context length, enabling it to handle extensive inputs effectively. Through rigorous enhancement processes such as supervised fine-tuning, proximal policy optimization, and direct preference optimization, Phi-3.5-mini ensures precise instruction adherence and robust safety measures. Key Features and Functionality: - Extended Context Handling: Supports up to 128K tokens, facilitating tasks that require processing long documents or conversations. - High-Quality Reasoning: Trained on reasoning-dense data to enhance problem-solving and analytical capabilities. - Efficient Performance: Delivers state-of-the-art results within a compact model size, making it suitable for resource-constrained environments. - Robust Safety Measures: Incorporates advanced optimization techniques to ensure safe and reliable outputs. Primary Value and User Solutions: Phi-3.5-mini addresses the need for a powerful yet efficient language model capable of handling extensive context lengths and complex reasoning tasks. Its compact size allows for deployment in environments with limited computational resources without compromising performance. By focusing on high-quality, reasoning-dense data, it provides users with accurate and contextually relevant outputs, making it ideal for applications in natural language understanding, content generation, and conversational AI.

Who Is the Company Behind Phi 3.5 mini?

  • Seller: Microsoft
  • Year Founded: 1975
  • HQ Location: Redmond, Washington
  • Twitter: @microsoft
    13,091,739 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    232,750 employees on LinkedIn®
  • Ownership: MSFT

Phi 3 mini 4k

The Phi-3 Mini-4K-Instruct is a lightweight, state-of-the-art language model developed by Microsoft, featuring 3.8 billion parameters. It is part of the Phi-3 model family and is designed to support a context length of 4,000 tokens. Trained on a combination of synthetic data and filtered publicly available websites, the model emphasizes high-quality, reasoning-dense content. Post-training enhancements, including supervised fine-tuning and direct preference optimization, have been applied to improve instruction adherence and safety measures. The Phi-3 Mini-4K-Instruct demonstrates robust performance across benchmarks assessing common sense, language understanding, mathematics, coding, long-context comprehension, and logical reasoning, positioning it as a leading model among those with fewer than 13 billion parameters. Key Features and Functionality: - Compact Architecture: With 3.8 billion parameters, the model offers a balance between performance and resource efficiency. - Extended Context Length: Supports processing of up to 4,000 tokens, enabling handling of longer inputs effectively. - High-Quality Training Data: Utilizes a curated dataset combining synthetic data and filtered web content, focusing on high-quality and reasoning-intensive information. - Enhanced Instruction Following: Post-training processes, including supervised fine-tuning and direct preference optimization, improve the model's ability to follow instructions accurately. - Versatile Performance: Excels in various tasks such as common sense reasoning, language understanding, mathematical problem-solving, coding, and logical reasoning. Primary Value and User Solutions: The Phi-3 Mini-4K-Instruct addresses the need for a powerful yet efficient language model suitable for environments with limited memory and computational resources. Its compact size and extended context capabilities make it ideal for applications requiring low latency and strong reasoning abilities. By delivering state-of-the-art performance in a resource-efficient package, it enables developers and researchers to integrate advanced language understanding and generation features into their applications without the overhead associated with larger models.

Who Is the Company Behind Phi 3 mini 4k?

  • Seller: Microsoft
  • Year Founded: 1975
  • HQ Location: Redmond, Washington
  • Twitter: @microsoft
    13,091,739 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    232,750 employees on LinkedIn®
  • Ownership: MSFT

Phi 3 small 128k

The Phi-3-Small-128K-Instruct is a 7-billion-parameter, state-of-the-art language model developed by Microsoft. It is part of the Phi-3 family and is designed to handle a context length of up to 128,000 tokens. Trained on a combination of synthetic data and filtered publicly available web content, the model emphasizes high-quality, reasoning-dense properties. Post-training processes, including supervised fine-tuning and direct preference optimization, have been applied to enhance its instruction-following capabilities and safety measures. The Phi-3-Small-128K-Instruct demonstrates robust performance across benchmarks testing common sense, language understanding, mathematics, coding, long-context comprehension, and logical reasoning, positioning it competitively among models of similar and larger sizes. Key Features and Functionality: - Extensive Context Handling: Supports a context length of up to 128,000 tokens, enabling the processing of long and complex inputs. - High-Quality Training Data: Utilizes a blend of synthetic and curated web data, focusing on content rich in reasoning and quality. - Advanced Post-Training Techniques: Incorporates supervised fine-tuning and direct preference optimization to improve instruction adherence and safety. - Versatile Performance: Excels in tasks requiring common sense, language understanding, mathematical reasoning, coding proficiency, and logical analysis. Primary Value and User Solutions: The Phi-3-Small-128K-Instruct model offers developers and researchers a powerful tool for building AI systems that require deep reasoning and the ability to process extensive contextual information. Its efficient architecture makes it suitable for memory and compute-constrained environments, while its strong performance in various reasoning tasks addresses the needs of applications demanding high levels of understanding and analysis. By providing a robust foundation for generative AI features, the model accelerates the development of advanced language and multimodal applications.

Who Is the Company Behind Phi 3 small 128k?

  • Seller: Microsoft
  • Year Founded: 1975
  • HQ Location: Redmond, Washington
  • Twitter: @microsoft
    13,091,739 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    232,750 employees on LinkedIn®
  • Ownership: MSFT

Phi 3 Small 8k

Smaller Phi-3 model variant with extended 8k token context and instruction capabilities.

Who Is the Company Behind Phi 3 Small 8k?

  • Seller: Microsoft
  • Year Founded: 1975
  • HQ Location: Redmond, Washington
  • Twitter: @microsoft
    13,091,739 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    232,750 employees on LinkedIn®
  • Ownership: MSFT

Phi 4 mini

The Phi-3 Mini-4K-Instruct is a lightweight, state-of-the-art language model developed by Microsoft, featuring 3.8 billion parameters. It is part of the Phi-3 model family and is designed to support a context length of 4,000 tokens. Trained on a combination of synthetic data and filtered publicly available websites, the model emphasizes high-quality, reasoning-dense content. Post-training enhancements, including supervised fine-tuning and direct preference optimization, have been applied to improve instruction adherence and safety measures. The Phi-3 Mini-4K-Instruct demonstrates robust performance across benchmarks assessing common sense, language understanding, mathematics, coding, long-context comprehension, and logical reasoning, positioning it as a leading model among those with fewer than 13 billion parameters. Key Features and Functionality: - Compact Architecture: With 3.8 billion parameters, the model offers a balance between performance and resource efficiency. - Extended Context Length: Supports processing of up to 4,000 tokens, enabling handling of longer inputs effectively. - High-Quality Training Data: Utilizes a curated dataset combining synthetic data and filtered web content, focusing on high-quality and reasoning-intensive information. - Enhanced Instruction Following: Post-training processes, including supervised fine-tuning and direct preference optimization, improve the model's ability to follow instructions accurately. - Versatile Performance: Excels in various tasks such as common sense reasoning, language understanding, mathematical problem-solving, coding, and logical reasoning. Primary Value and User Solutions: The Phi-3 Mini-4K-Instruct addresses the need for a powerful yet efficient language model suitable for environments with limited memory and computational resources. Its compact size and extended context capabilities make it ideal for applications requiring low latency and strong reasoning abilities. By delivering state-of-the-art performance in a resource-efficient package, it enables developers and researchers to integrate advanced language understanding and generation features into their applications without the overhead associated with larger models.

Who Is the Company Behind Phi 4 mini?

  • Seller: Microsoft
  • Year Founded: 1975
  • HQ Location: Redmond, Washington
  • Twitter: @microsoft
    13,091,739 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    232,750 employees on LinkedIn®
  • Ownership: MSFT

Phi 4 mini reasoning

Phi-4-mini-reasoning is a compact, transformer-based language model developed by Microsoft, specifically optimized for mathematical reasoning tasks. With 3.8 billion parameters and support for a 128K token context length, it delivers high-quality, step-by-step problem-solving capabilities in environments where computational resources or latency are constrained. Fine-tuned using synthetic mathematical data generated by a more advanced model, Phi-4-mini-reasoning excels in multi-step, logic-intensive problem-solving scenarios, making it suitable for applications such as formal proof generation, symbolic computation, and advanced word problems. Key Features and Functionality: - Optimized for Mathematical Reasoning: Designed to handle complex, multi-step mathematical problems with structured logic and analytical thinking. - Compact Architecture: Balances reasoning ability with efficiency, enabling deployment in resource-constrained environments. - Extended Context Length: Supports up to 128K tokens, allowing for comprehensive context retention across problem-solving steps. - Fine-Tuned with Synthetic Data: Trained on a diverse set of over one million math problems, enhancing its reasoning performance. Primary Value and Problem Solving: Phi-4-mini-reasoning addresses the need for efficient, high-quality mathematical reasoning in scenarios where computational resources are limited. Its compact size and optimized performance make it ideal for educational applications, embedded tutoring systems, and deployments on edge or mobile devices. By maintaining context across multiple steps and applying structured logic, it provides accurate and reliable solutions for complex mathematical problems, thereby enhancing learning experiences and supporting advanced analytical tasks.

Who Is the Company Behind Phi 4 mini reasoning?

  • Seller: Microsoft
  • Year Founded: 1975
  • HQ Location: Redmond, Washington
  • Twitter: @microsoft
    13,091,739 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    232,750 employees on LinkedIn®
  • Ownership: MSFT

StableLM 2 1.6b

StableLM 2 1.6B is a 1.6 billion parameter decoder-only language model developed by Stability AI. It is pre-trained on 2 trillion tokens from diverse multilingual and code datasets over two epochs. The model is designed to generate coherent and contextually relevant text, making it suitable for a wide range of natural language processing tasks. Key Features and Functionality: - Transformer Decoder Architecture: StableLM 2 1.6B utilizes a decoder-only transformer architecture, similar to LLaMA, with specific modifications to enhance performance. - Rotary Position Embeddings: Incorporates Rotary Position Embeddings applied to the first 25% of head embedding dimensions, improving throughput. - Layer Normalization: Employs LayerNorm with learned bias terms, differing from RMSNorm, to stabilize training and improve convergence. - Bias Configuration: Removes all bias terms from feed-forward networks and multi-head self-attention layers, except for the biases of the query, key, and value projections, optimizing computational efficiency. - Advanced Tokenization: Utilizes the Arcade100k tokenizer, a BPE tokenizer extended from OpenAI's tiktoken.cl100k_base, with digit splitting into individual tokens to enhance numerical understanding. Primary Value and User Solutions: StableLM 2 1.6B offers a robust solution for developers and researchers seeking a powerful language model capable of generating high-quality text across various applications. Its extensive pre-training on diverse datasets ensures versatility in handling multiple languages and code, making it ideal for tasks such as content creation, code generation, and multilingual translation. The model's architecture and training methodologies provide a balance between performance and computational efficiency, addressing the need for scalable and effective language models in the AI community.

Who Is the Company Behind StableLM 2 1.6b?

  • Seller: Stability AI
  • HQ Location: London
  • Twitter: @StabilityAI
    256,849 Twitter followers
  • LinkedIn® Page: www.linkedin.com
    189 employees on LinkedIn®

step-1 8k

Step-1 8k is a large-scale language model developed by StepFun, designed to understand and generate natural language text across various domains. With a context length of 8,000 tokens, it can process substantial input and output, making it suitable for tasks such as content creation, multilingual communication, question answering, and logical reasoning. Additionally, Step-1 8k exhibits strong mathematical and coding capabilities, supporting applications in scientific computation and software development. Key Features and Functionality: - Extensive Context Processing: Handles up to 8,000 tokens, allowing for comprehensive understanding and generation of lengthy texts. - Versatile Language Tasks: Excels in content generation, translation, summarization, and conversational AI. - Mathematical and Coding Proficiency: Capable of performing complex calculations and generating code snippets, aiding in scientific and programming tasks. - High Cost-Performance Ratio: Offers a balance between performance and cost, making it accessible for various applications. Primary Value and User Solutions: Step-1 8k enhances productivity by automating and streamlining language-related tasks. Its ability to process extensive context ensures coherent and contextually relevant outputs, benefiting professionals in content creation, software development, and data analysis. By integrating Step-1 8k, users can achieve efficient and accurate results in their respective fields.

Who Is the Company Behind step-1 8k?

Sutra

Multilingual Mixture-of-Experts model supporting 50+ languages with better MMLU performance and reduced hallucinations using online knowledge.

Who Is the Company Behind Sutra?

  • Seller: Two AI
  • Year Founded: 2021
  • HQ Location: Silicon Valley, US
  • LinkedIn® Page: www.linkedin.com
    49 employees on LinkedIn®
Jeffrey Lin
JL
Researched and written by Jeffrey Lin
Updated April 9, 2026