Abstract
We present Jamba, a new base large language model based on a novel hybrid Transformer-Mamba mixture-of-experts (MoE) architecture. Specifically, Jamba interleaves blocks of Transformer and Mamba layers, enjoying the benefits of both model families. MoE is added in some of these layers to increase model capacity while keeping active parameter usage manageable. This flexible architecture allows resource- and objective-specific configurations. In the particular configuration we have implemented, we end up with a powerful model that fits in a single 80GB GPU. Built at large scale, Jamba provides high throughput and small memory footprint compared to vanilla Transformers, and at the same time state-of-the-art performance on standard language model benchmarks and long-context evaluations. Remarkably, the model presents strong results for up to 256K tokens context length. We study various architectural decisions, such as how to combine Transformer and Mamba layers, and how to mix experts, and show that some of them are crucial in large scale modeling. We also describe several interesting properties of these architectures which the training and evaluation of Jamba have revealed, and plan to release checkpoints from various ablation runs, to encourage further exploration of this novel architecture. We make the weights of our implementation of Jamba publicly available under a permissive license.
Community
good
TechxGenus (Hao Jiang) made a 69M param model https://huggingface.co/TechxGenus/Mini-Jamba
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models (2024)
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference (2024)
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference (2024)
- ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching (2024)
- Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks (2024)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment:
@librarian-bot
recommend
Jamba: Revolutionizing Language Models with a Hybrid Transformer Approach
Links π:
π Subscribe: https://www.youtube.com/@Arxflix
π Twitter: https://x.com/arxflix
π LMNT (Partner): https://lmnt.com/
Models citing this paper 4
Datasets citing this paper 0
No dataset linking this paper