References

Every paper, book, and software citation in this book, sorted alphabetically by first-author surname. Click an entry’s anchor to share a deep link, or follow the arXiv / DOI / URL for the source.

  1. [almohy2011computing] Awad H. Al-Mohy and Nicholas J. Higham (2011). Computing the Action of the Matrix Exponential, with an Application to Exponential Integrators. SIAM Journal on Scientific Computing. Vol. 33, no. 2. pp. 488-511.
  2. [antoulas2005approximation] Athanasios C. Antoulas (2005). Approximation of Large-Scale Dynamical Systems. SIAM.
  3. [arora2024zoology] Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Zhang, James Zou, Atri Rudra, and Christopher Ré (2024). Zoology: Measuring and Improving Recall in Efficient Language Models. International Conference on Learning Representations (ICLR). link.
  4. [arora2024simple] Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré (2024). Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff. arXiv preprint. link.
  5. [babaei2025walrus] Hossein Babaei, Mel White, and Richard G. Baraniuk (2025). W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling. arXiv preprint. link.
  6. [bai2024longbench] Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, and others (2024). LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. Annual Meeting of the Association for Computational Linguistics (ACL). link.
  7. [bambateam2024bamba] Bamba Team (IBM, Princeton, CMU, UIUC) (2024). Bamba: Inference-Efficient Hybrid Mamba2 Model. Hugging Face blog, https://huggingface.co/blog/bamba.
  8. [basu2026content] Abhinaba Basu (2026). When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models. arXiv preprint. link.
  9. [beck2024xlstm] Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter (2024). xLSTM: Extended Long Short-Term Memory. Advances in Neural Information Processing Systems (NeurIPS). link.
  10. [behrouz2025titans] Ali Behrouz, Peilin Zhong, and Vahab Mirrokni (2025). Titans: Learning to Memorize at Test Time. arXiv preprint. link.
  11. [benettin1980lyapunov] Giancarlo Benettin, Luigi Galgani, Antonio Giorgilli, and Jean-Marie Strelcyn (1980). Lyapunov characteristic exponents for smooth dynamical systems and for Hamiltonian systems; a method for computing all of them. Meccanica. Vol. 15, no. 1. pp. 9-20.
  12. [bick2025understanding] Aviv Bick, Eric P. Xing, and Albert Gu (2025). Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism. arXiv preprint. link.
  13. [blelloch1990prefix] Guy E. Blelloch (1990). Prefix sums and their applications. Synthesis of Parallel Algorithms. pp. 35-60.
  14. [choromanski2021performer] Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller (2021). Rethinking Attention with Performers. International Conference on Learning Representations (ICLR). link.
  15. [dahlquist1956convergence] Germund Dahlquist (1956). Convergence and stability in the numerical integration of ordinary differential equations. Mathematica Scandinavica. Vol. 4. pp. 33-53.
  16. [dao2024mamba2] Tri Dao and Albert Gu (2024). Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. International Conference on Machine Learning (ICML). link.
  17. [de2024griffin] Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, and Caglar Gulcehre (2024). Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models. arXiv preprint. link.
  18. [dong2024hymba] Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, and others (2024). Hymba: A Hybrid-head Architecture for Small Language Models. arXiv preprint. link.
  19. [golub2013matrix] Gene H. Golub and Charles F. Van Loan (2013). Matrix Computations. Johns Hopkins University Press.
  20. [gu2020hippo] Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré (2020). HiPPO: Recurrent Memory with Optimal Polynomial Projections. Advances in Neural Information Processing Systems (NeurIPS). link.
  21. [gu2022s4] Albert Gu, Karan Goel, and Christopher Ré (2022). Efficiently Modeling Long Sequences with Structured State Spaces. International Conference on Learning Representations (ICLR). link.
  22. [gu2022s4d] Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré (2022). On the Parameterization and Initialization of Diagonal State Space Models. Advances in Neural Information Processing Systems (NeurIPS). link.
  23. [gu2023train] Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré (2023). How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections. International Conference on Learning Representations (ICLR). link.
  24. [gu2024mamba] Albert Gu and Tri Dao (2024). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. Conference on Language Modeling (COLM). link.
  25. [gupta2022diagonal] Ankit Gupta, Albert Gu, and Jonathan Berant (2022). Diagonal State Spaces are as Effective as Structured State Spaces. Advances in Neural Information Processing Systems (NeurIPS). link.
  26. [hairer1993ordinary] Ernst Hairer, Syvert P. Nørsett, and Gerhard Wanner (1993). Solving Ordinary Differential Equations I: Nonstiff Problems. Springer. Vol. 8.
  27. [hairer1996ordinary] Ernst Hairer and Gerhard Wanner (1996). Solving Ordinary Differential Equations II: Stiff and Differential-Algebraic Problems. Springer. Vol. 14.
  28. [hairer2006geometric] Ernst Hairer, Christian Lubich, and Gerhard Wanner (2006). Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations. Springer. Vol. 31.
  29. [anonymous2025lyapunov] John Timothy Halloran, Manbir S. Gulati, and Paul F. Roysdon (2025). Mamba State-Space Models Are Lyapunov-Stable Learners. Transactions on Machine Learning Research (TMLR). link.
  30. [hochbruck2010exponential] Marlis Hochbruck and Alexander Ostermann (2010). Exponential integrators. Acta Numerica. Vol. 19. pp. 209-286.
  31. [hsieh2024ruler] Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg (2024). RULER: What's the Real Context Size of Your Long-Context Language Models?. First Conference on Language Modeling (COLM). link.
  32. [jelassi2024repeat] Samy Jelassi, David Brandfonbrener, Sham M. Kakade, and Eran Malach (2024). Repeat After Me: Transformers are Better than State Space Models at Copying. arXiv preprint. link.
  33. [katharopoulos2020transformers] Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret (2020). Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. International Conference on Machine Learning (ICML). link.
  34. [kimiteam2025kimi] Kimi Team (2025). Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv preprint. link.
  35. [kokotovic1986singular] Petar Kokotović, Hassan K. Khalil, and John O'Reilly (1986). Singular Perturbation Methods in Control: Analysis and Design. Academic Press.
  36. [lahoti2026mamba3] Aakash Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, and Albert Gu (2026). Mamba-3: Improved Sequence Modeling using State Space Principles. International Conference on Learning Representations (ICLR). link.
  37. [lax1956survey] Peter D. Lax and Robert D. Richtmyer (1956). Survey of the stability of linear finite difference equations. Communications on Pure and Applied Mathematics. Vol. 9, no. 2. pp. 267-293.
  38. [lee2025understanding] Hyunji Lee, Wenhao Yu, Hongming Zhang, Kaixin Ma, Jiyeon Kim, Dong Yu, and Minjoon Seo (2025). Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling. arXiv preprint. link.
  39. [lieber2024jamba] Opher Lieber, Barak Lenz, Hofit Bata, and others (2024). Jamba: A Hybrid Transformer-Mamba Language Model. arXiv preprint. link.
  40. [liu2024longhorn] Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, and Qiang Liu (2024). Longhorn: State Space Models are Amortized Online Learners. arXiv preprint. link.
  41. [merrill2023parallelism] William Merrill and Ashish Sabharwal (2023). The Parallelism Tradeoff: Limitations of Log-Precision Transformers. Transactions of the Association for Computational Linguistics (TACL). link.
  42. [merrill2024illusion] William Merrill, Jackson Petty, and Ashish Sabharwal (2024). The Illusion of State in State-Space Models. Proceedings of the 41st International Conference on Machine Learning (ICML). link.
  43. [nvidia2025nemotron] NVIDIA (2025). Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models. arXiv preprint. link.
  44. [olsson2022incontext] Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, and others (2022). In-context Learning and Induction Heads. arXiv preprint. link.
  45. [park2024numerical] Jaesung R. Park, Jaewook J. Suh, Youngjoon Hong, and Ernest K. Ryu (2024). Numerical Analysis of HiPPO-LegS ODE for Deep State Space Models. arXiv preprint. link.
  46. [peng2025rwkv] Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, William Merrill, and others (2025). RWKV-7 "Goose" with Expressive Dynamic State Evolution. arXiv preprint. link.
  47. [poli2023hyena] Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré (2023). Hyena Hierarchy: Towards Larger Convolutional Language Models. International Conference on Machine Learning (ICML). link.
  48. [poli2024mechanistic] Michael Poli, Armin W. Thomas, Eric Nguyen, and others (2024). Mechanistic Design and Scaling of Hybrid Architectures. International Conference on Machine Learning (ICML). link.
  49. [qin2024hgrn2] Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong (2024). HGRN2: Gated Linear RNNs with State Expansion. Conference on Language Modeling (COLM). link.
  50. [ren2024samba] Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu Chen (2024). Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling. arXiv preprint. link.
  51. [ren2025decoder] Liliang Ren, Congcong Chen, Haoran Xu, Young Jin Kim, and others (2025). Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation. arXiv preprint. link.
  52. [schlag2021linear] Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber (2021). Linear Transformers Are Secretly Fast Weight Programmers. International Conference on Machine Learning (ICML). link.
  53. [seif2022impact] Alireza Seif, Sarah A. M. Loos, Gennaro Tucci, Édgar Roldán, and Sebastian Goldt (2022). The impact of memory on learning sequence-to-sequence tasks. arXiv preprint. link.
  54. [shaham2022scrolls] Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, and others (2022). SCROLLS: Standardized CompaRison Over Long Language Sequences. Conference on Empirical Methods in Natural Language Processing (EMNLP). link.
  55. [smith2023simplified] Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman (2023). Simplified State Space Layers for Sequence Modeling. International Conference on Learning Representations (ICLR). link.
  56. [sun2023retnet] Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei (2023). Retentive Network: A Successor to Transformer for Large Language Models. arXiv preprint. link.
  57. [tay2021long] Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler (2021). Long Range Arena: A Benchmark for Efficient Transformers. International Conference on Learning Representations (ICLR). link.
  58. [hunyuanteam2025hunyuan] Tencent Hunyuan Team (2025). Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought. arXiv preprint. link.
  59. [trefethen1997numerical] Lloyd N. Trefethen and David Bau (1997). Numerical Linear Algebra. SIAM.
  60. [tustin1947method] Arnold Tustin (1947). A method of analysing the behaviour of linear systems in terms of time series. Journal of the Institution of Electrical Engineers. Vol. 94. pp. 130-142.
  61. [vanloan1978computing] Charles F. Van Loan (1978). Computing integrals involving the matrix exponential. IEEE Transactions on Automatic Control. Vol. 23, no. 3. pp. 395-404.
  62. [widrow1960adaptive] Bernard Widrow and Marcian E. Hoff (1960). Adaptive Switching Circuits. 1960 IRE WESCON Convention Record, Part 4. pp. 96-104.
  63. [yang2024gla] Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim (2024). Gated Linear Attention Transformers with Hardware-Efficient Training. International Conference on Machine Learning (ICML). link.
  64. [yang2024deltanet] Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim (2024). Parallelizing Linear Transformers with the Delta Rule over Sequence Length. Advances in Neural Information Processing Systems (NeurIPS). link.
  65. [yang2025gateddeltanet] Songlin Yang, Jan Kautz, and Ali Hatamizadeh (2025). Gated Delta Networks: Improving Mamba2 with Delta Rule. International Conference on Learning Representations (ICLR). link.
  66. [yu2023robustifying] Annan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney, and N. Benjamin Erichson (2023). Robustifying State-space Models for Long Sequences via Approximate Diagonalization. arXiv preprint. link.
  67. [zheng2026amor] Haoran Zheng and Chen Shani (2026). When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models. arXiv preprint. link.