Artificial intelligence architecture is no longer adequately described by model size alone. Dense Transformers remain the reference architecture for language and multimodal reasoning, but production systems increasingly combine conditional computation, retrieval, memory, tools, verifiers, edge-cloud routing, observability and governance. This review makes three engineering claims. First, sparse Mixture-of-Experts models are currently the clearest capacity-scaling pattern, because they decouple total parameters from active per-token computation, although routing imbalance and distributed communication remain hard constraints. Second, state-space, recurrent and linear attention hybrids are best interpreted as attention-budgeting architectures: they reduce KV-cache and long-context costs, but do not yet displace dense attention in every reasoning regime. Third, JEPA-style latent world models change the learning objective from surface-token or pixel prediction to representation prediction, which is strategically important for perception and planning but still not a drop-in replacement for general language interfaces. To make the maturity claims auditable, this review uses a PRISMA-inspired search protocol, an explicit technology readiness rubric, quantitative comparison tables, hardware and memory-bandwidth analysis, deployment and reproducibility categories, and failure cases for RAG and agents. The main conclusion is that the optimal architecture is task- and constraint-dependent: small dense or hybrid models are often preferred for real-time edge inference, RAG and graph memory for changing enterprise knowledge, frontier dense or sparse models for difficult reasoning, and agentic workflows only when tool permissions, rollback, provenance and human oversight are engineered as first-class components.

Current Trends in Artificial Intelligence Architectures: From Model Scaling to System Intelligence, Post-Transformer Hybrids and World Models

Rampone, Salvatore
2026-01-01

Abstract

Artificial intelligence architecture is no longer adequately described by model size alone. Dense Transformers remain the reference architecture for language and multimodal reasoning, but production systems increasingly combine conditional computation, retrieval, memory, tools, verifiers, edge-cloud routing, observability and governance. This review makes three engineering claims. First, sparse Mixture-of-Experts models are currently the clearest capacity-scaling pattern, because they decouple total parameters from active per-token computation, although routing imbalance and distributed communication remain hard constraints. Second, state-space, recurrent and linear attention hybrids are best interpreted as attention-budgeting architectures: they reduce KV-cache and long-context costs, but do not yet displace dense attention in every reasoning regime. Third, JEPA-style latent world models change the learning objective from surface-token or pixel prediction to representation prediction, which is strategically important for perception and planning but still not a drop-in replacement for general language interfaces. To make the maturity claims auditable, this review uses a PRISMA-inspired search protocol, an explicit technology readiness rubric, quantitative comparison tables, hardware and memory-bandwidth analysis, deployment and reproducibility categories, and failure cases for RAG and agents. The main conclusion is that the optimal architecture is task- and constraint-dependent: small dense or hybrid models are often preferred for real-time edge inference, RAG and graph memory for changing enterprise knowledge, frontier dense or sparse models for difficult reasoning, and agentic workflows only when tool permissions, rollback, provenance and human oversight are engineered as first-class components.
2026
artificial intelligence; foundation models; post-Transformer architectures; mixture of experts; state-space models; JEPA; world models; retrieval-augmented generation; agentic AI; multimodal AI; liquid neural networks; edge AI; neuro-symbolic AI; AI governance; technology readiness
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12070/76365
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact