arxiv_ml 95% Match Research Paper LLM Researchers,Generative AI Developers,NLP Engineers,AI Model Architects 2 weeks ago

Large Language Diffusion Models

large-language-models › model-architecture

📄 Abstract

Abstract: The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model trained from scratch under the pre-training and supervised fine-tuning (SFT) paradigm. LLaDA employs a forward data masking process and a reverse generation process, parameterized by a Transformer to predict masked tokens. It provides a principled generative approach for probabilistic inference by optimizing a likelihood lower bound. Across extensive benchmarks on general tasks, math, code, and so on, LLaDA demonstrates strong scalability and performs comparably to our self-constructed ARM baselines. Remarkably, LLaDA 8B is competitive with strong LLMs like LLaMA3 8B in in-context learning and, after SFT, exhibits impressive instruction-following abilities in case studies such as multi-turn dialogue. Moreover, LLaDA addresses the reversal curse, surpassing GPT-4o in a reversal poem completion task. Our findings show the promise of diffusion models for language modeling at scale and challenge the common assumption that core LLM capabilities discussed above inherently depend on ARMs. Project page and codes: https://ml-gsai.github.io/LLaDA-demo/.

Authors (10)

Shen Nie

Fengqi Zhu

Zebin You

Xiaolu Zhang

Jingyang Ou

Jun Hu

+4 more

Submitted

February 14, 2025

arXiv Category

cs.CL

arXiv PDF

Key Contributions

Introduces LLaDA, a diffusion model trained from scratch for language generation, challenging the dominance of autoregressive models. LLaDA demonstrates strong scalability and competitive performance with ARMs across various tasks, exhibits impressive instruction-following abilities after SFT, and addresses the 'reversal curse'.

Business Value

Offers a new paradigm for building powerful LLMs, potentially leading to more efficient training, novel generation capabilities, and improved performance in specific tasks like code generation or complex instruction following.

Paper Metadata

Innovation Type

Model Architecture/Training Method

Deployment Feasibility

Moderate, requires significant computational resources for training diffusion models at scale.

Limitations Addressed

Dominance of autoregressive models in LLMs,Potential limitations of autoregressive generation (e.g., reversal curse),Scalability challenges in LLM training

Technical Tags

diffusion modelslarge language models (LLMs)autoregressive models (ARMs)Transformergenerative modelspre-trainingsupervised fine-tuning (SFT)scalabilityinstruction followingreversal curse

Research Topics

Generative AILarge Language ModelsDiffusion ModelsModel ArchitecturesNatural Language Generation

Methods & Architectures

Diffusion Model TrainingForward Data MaskingReverse Generation ProcessTransformer ParameterizationPre-trainingSupervised Fine-Tuning (SFT) Diffusion ModelsTransformersAutoregressive Models (ARMs)

Applications & Tasks

Natural Language Generation Text Generation AI Model Development Developing non-autoregressive LLMsImproving LLM scalabilityAddressing the reversal curse in generation Training diffusion models for language generation from scratchAchieving comparable performance to ARMsDemonstrating instruction-following abilities

Datasets & Benchmarks

Benchmarks

general tasks • math • code

Related Fields

Generative AIDiffusion ModelsLarge Language ModelsNatural Language ProcessingDeep Learning

Keywords

diffusion modelsLLMsautoregressiveTransformergenerative AIpre-trainingSFTscalabilityinstruction followingreversal curseLLaDA

Academic Context

#Generative AI#Large Language Models#Diffusion Models#Model Architectures#Natural Language Generation

Companies & Organizations

Companies Mentioned

Meta AI (LLaMA)

Commercial Potential

Potential Products

Next-generation LLM architecturesSpecialized text generation modelsTools for efficient LLM training

Target Industries

TechnologySoftware DevelopmentContent CreationResearch

Use Case Examples

Generating code or complex structured text more efficiently.Developing LLMs with improved capabilities in tasks prone to the 'reversal curse'.Building large-scale language models with potentially different training dynamics.

Competitive Edge

Challenges the established autoregressive paradigm for LLMs by demonstrating the viability and effectiveness of diffusion models in this domain.

Market Opportunity

Massive and rapidly growing market for LLMs.

Revenue Models

Model licensingAPI accessspecialized applications.

Resource Requirements

Compute Needs

Very high, for training diffusion models from scratch at scale.

Data Requirements

Large text corpora for pre-training.

Deployment Constraints

Computational cost and complexity of diffusion model inference compared to autoregressive models.

Scalability

Demonstrates strong scalability, comparable to ARMs.

Production Readiness

Maturity Level

Research

Time to Market

Long

Patent Potential

Moderate

View Full Paper Back to Papers