Iqra Journal of Engineering and Computing

A Comparative Review of Transformer Architectures: Evolution, Efficiency, Applications, and Future Directions

Research Article 2
- Volume 2, Issue 1 2026
By Israr Ali,Ali Ahmed Siddiqui,Aarij Mahmood Hussaan
10.21621/ijec.20260201.05
Keywords: Transformer, self-attention, BERT, GPT, large language models, multimodal learning, Mixture-of-Experts, efficient attention, Vision Transformer, state-space models, benchmark evaluation, long-context reasoning, open-weight models

Transformer architectures have become a central foundation of modern artificial intelligence because of their scalability, parallel computation, and representation learning across language, vision, code, and multimodal domains. Since the original Transformer, the architecture has evolved into encoder-only, decoder-only, encoder-decoder, efficient attention, vision, multimodal, sparse Mixture-of-Experts, open-weight, state-space, recurrent, and hybrid model families. This review presents a structured and benchmark-informed comparison of Transformer architectures using a defined literature selection protocol, inclusion and exclusion criteria, model-family coding dimensions, and source-verification rules. The paper compares major architectures with respect to training objective, attention or sequence-modeling mechanism, computational complexity, scalability, context length, modality support, openness, benchmark behavior, reproducibility, deployment suitability, and architectural trade-offs. Representative benchmarks, including GLUE/SuperGLUE, MMLU/MMLU-Pro, HELM, HumanEval, SWE-bench, LongBench, RULER, ImageNet, and MMMU, are used to support critical discussion rather than relying only on descriptive model summaries. Recent developments, including GPT-4o, Gemini 2.5, Llama 4, DeepSeek-R1, Qwen3, Jamba, RecurrentGemma/Griffin, Hyena, Mamba, and xLSTM, are incorporated to reflect the 2024-2025 research landscape. The review identifies evidence-backed gaps in long-context reasoning, multimodal evaluation, cost-aware comparison, open versus closed model reproducibility, safety, interpretability, and domain-specific adaptation. Overall, the study provides a rigorous comparative foundation for researchers and practitioners evaluating Transformer-based and hybrid artificial intelligence systems.

Submission Date: 2 Jun, 2026 Reviews Completed: 18 Jun, 2026
Acceptance Date: 20 Jun, 2026 Publication Date: 23 Jun, 2026

Share this paper


Want to publish in ?
Send us your paper for review
29
Authors
15
Research Papers
0
Citations