Introduce an entity-centric Trimodal Structured-Data Benchmark (TSDBench) for multimodal learning.
TSDBench validates multimodal information gains via ablation and partial information decomposition (PID).
Results reveal cross-modal gains are task-dependent, asymmetric, and vary with prediction targets.
Flow-PID exposes gaps between theoretical multimodal synergy and what current models can exploit.
| [1] | Dua D. and Graff C. (2019). UCI Machine Learning Repository. |
| [2] | Vanschoren J., van Rijn J. N., Bischl B., et al. (2013). OpenML: networked science in machine learning. ACM SIGKDD Explor. Newsl. 15:49−60. DOI:10.1145/2641190.2641198 |
| [3] | Chen T. and Guestrin C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Association for Computing Machinery) 785−794. DOI:10.1145/2939672.2939785 |
| [4] | Prokhorenkova L., Gusev G., Vorobev A., et al. (2018). CatBoost: unbiased boosting with categorical features. Bengio S., Wallach H., Larochelle H., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)31: 6638–6648. DOI: 10.48550/arXiv.1706.09516 |
| [5] | Gorishniy Y., Rubachev I., Khrulkov V., et al. (2021). Revisiting Deep Learning Models for Tabular Data. Ranzato M., Beygelzimer A., Dauphin Y., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)34: 18932–18943. DOI: 10.48550/arXiv.2106.11959 |
| [6] | Hollmann N., Müller S., Eggensperger K., et al. (2023). TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second. International conference on learning representations. DOI:10.48550/arXiv.2207.01848 |
| [7] | Grinsztajn L., Oyallon E. and Varoquaux G. (2022). Why do tree-based models still outperform deep learning on typical tabular data? Koyejo S., Mohamed S., Agarwal A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)35: 507–520. DOI: 10.52202/068431-0037 |
| [8] | Makridakis S., Spiliotis E. and Assimakopoulos V. (2020). The M4 Competition: 100, 000 time series and 61 forecasting methods. Int. J. Forecast. 36:54−74. DOI:10.1016/j.ijforecast.2019.04.014 |
| [9] | Godahewa R., Bergmeir C., Webb G. I., et al. (2021). Monash Time Series Forecasting Archive. Vanschoren J. and Yeung S. (ed). Proceedings of the neural information processing systems track on datasets and benchmarks 1. DOI: 10.48550/arXiv.2105.06643 |
| [10] | Zhou H., Zhang S., Peng J., et al. (2021). Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 35:11106−11115. DOI:10.1609/aaai.v35i12.17325 |
| [11] | Wu H., Xu J., Wang J., et al. (2021). Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. Ranzato M., Beygelzimer A., Dauphin Y., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)34: 22419–22430. DOI: 10.48550/arXiv.2106.13008 |
| [12] | Zeng A., Chen M., Zhang L., et al. (2023). Are Transformers Effective for Time Series Forecasting? Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 37:11121−11128. DOI:10.1609/aaai.v37i9.26317 |
| [13] | Nie Y., Nguyen N. H., Sinthong P., et al. (2023). A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. International conference on learning representations. DOI:10.48550/arXiv.2211.14730 |
| [14] | Liu Y., Hu T., Zhang H., et al. (2024). iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. Kim B., Yue Y., Chaudhuri S., et al. (ed). International conference on learning representations 2024: 11116–11140. DOI: 10.48550/arXiv.2310.06625 |
| [15] | Kipf T. N. and Welling M. (2017). Semi-Supervised Classification with Graph Convolutional Networks. International conference on learning representations. DOI:10.48550/arXiv.1609.02907 |
| [16] | Hamilton W. L., Ying R. and Leskovec J. (2017). Inductive Representation Learning on Large Graphs. Guyon I., Luxburg U. V., Bengio S., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)30: 1024–1034. DOI: 10.48550/arXiv.1706.02216 |
| [17] | Veličković P., Cucurull G., Casanova A., et al. (2018). Graph Attention Networks. International conference on learning representations. DOI:10.48550/arXiv.1710.10903 |
| [18] | Xu K., Hu W., Leskovec J., et al. (2019). How Powerful are Graph Neural Networks? International conference on learning representations. DOI:10.48550/arXiv.1810.00826 |
| [19] | Hu W., Fey M., Zitnik M., et al. (2020). Open Graph Benchmark: Datasets for Machine Learning on Graphs. Larochelle H., Ranzato M., Hadsell R., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)33: 22118–22133. DOI: 10.48550/arXiv.2005.00687 |
| [20] | Huang S., Poursafaei F., Danovitch J., et al. (2023). Temporal Graph Benchmark for Machine Learning on Temporal Graphs. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 2056–2073. DOI: 10.52202/075280-0099 |
| [21] | Bazhenov G., Platonov O. and Prokhorenkova L. (2025). GraphLand: Evaluating Graph Machine Learning Models on Diverse Industrial Data. Belgrave D., Zhang C., Lin H., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)38: 110211–110237. DOI: 10.52202/085713-3321 |
| [22] | Fu Y., Shao Z., Yu C., et al. (2026). Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis. DOI: 10.48550/arXiv.2607.01918 |
| [23] | Yu C., Wang F., Shao Z., et al. (2025). GinAR+: A Robust End-to-End Framework for Multivariate Time Series Forecasting With Missing Values. IEEE Trans. Knowl. Data Eng. 37:4635−4648. DOI:10.1109/tkde.2025.3569649 |
| [24] | Shao Z., Wang F., Xu Y., et al. (2025). Exploring Progress in Multivariate Time Series Forecasting: Comprehensive Benchmarking and Heterogeneity Analysis. IEEE Trans. Knowl. Data Eng. 37:291−305. DOI:10.1109/tkde.2024.3484454 |
| [25] | Li Y., Yu R., Shahabi C., et al. (2018). Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. International conference on learning representations. DOI:10.48550/arXiv.1707.01926 |
| [26] | Yu B., Yin H. and Zhu Z. (2018). Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. Proceedings of the twenty-seventh international joint conference on artificial intelligence (International Joint Conferences on Artificial Intelligence Organization) 3634−3640. DOI:10.24963/ijcai.2018/505 |
| [27] | Wu Z., Pan S., Long G., et al. (2019). Graph WaveNet for Deep Spatial-Temporal Graph Modeling. Proceedings of the twenty-eighth international joint conference on artificial intelligence (International Joint Conferences on Artificial Intelligence Organization) 1907−1913. DOI:10.24963/ijcai.2019/264 |
| [28] | Bai L., Yao L., Li C., et al. (2020). Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting. Larochelle H., Ranzato M., Hadsell R., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)33: 17804–17815. DOI: 10.48550/arXiv.2007.02842 |
| [29] | Liu X., Xia Y., Liang Y., et al. (2023). LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 75354–75371. DOI: 10.52202/075280-3293 |
| [30] | Fey M., Hu W., Huang K., et al. (2024). Position: Relational Deep Learning - Graph Representation Learning on Relational Databases. Salakhutdinov R., Kolter Z., Heller K., et al. (ed). Proceedings of the 41st International Conference on Machine Learning (PMLR)235: 13592–13607. DOI: 10.48550/arXiv.2312.04615 |
| [31] | Robinson J., Ranjan R., Hu W., et al. (2024). RelBench: A Benchmark for Deep Learning on Relational Databases. Globerson A., Mackey L., Belgrave D., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)37: 21330–21341. DOI: 10.52202/079017-0672 |
| [32] | Hu W., Yuan Y., Zhang Z., et al. (2024). PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning. DOI: 10.48550/arXiv.2404.00776 |
| [33] | Baltrušaitis T., Ahuja C. and Morency L. P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41:423−443. DOI:10.1109/tpami.2018.2798607 |
| [34] | Zadeh A., Chen M., Poria S., et al. (2017). Tensor Fusion Network for Multimodal Sentiment Analysis. Proceedings of the 2017 conference on empirical methods in natural language processing (Association for Computational Linguistics) 1103−1114. DOI:10.18653/v1/d17-1115 |
| [35] | Liu Z., Shen Y., Lakshminarasimhan V. B., et al. (2018). Efficient Low-rank Multimodal Fusion With Modality-Specific Factors. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Association for Computational Linguistics) 2247−2256. DOI:10.18653/v1/p18-1209 |
| [36] | Tsai Y. H. H., Bai S., Liang P. P., et al. (2019). Multimodal Transformer for Unaligned Multimodal Language Sequences. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Association for Computational Linguistics) 6558−6569. DOI:10.18653/v1/p19-1656 |
| [37] | Jaegle A., Gimeno F., Brock A., et al. (2021). Perceiver: General Perception with Iterative Attention. Meila M. and Zhang T. (ed). Proceedings of the 38th International Conference on Machine Learning (PMLR)139: 4651–4664. DOI: 10.48550/arXiv.2103.03206 |
| [38] | Liang P. P., Lyu Y., Fan X., et al. (2021). MultiBench: Multiscale Benchmarks for Multimodal Representation Learning. Vanschoren J. and Yeung S. (ed). Proceedings of the neural information processing systems track on datasets and benchmarks 1. DOI: 10.48550/arXiv.2107.07502 |
| [39] | Williams P. L. and Beer R. D. (2010). Nonnegative Decomposition of Multivariate Information. DOI: 10.48550/arXiv.1004.2515 |
| [40] | Bertschinger N., Rauh J., Olbrich E., et al. (2014). Quantifying Unique Information. Entropy 16:2161−2183. DOI:10.3390/e16042161 |
| [41] | Liang P. P., Cheng Y., Fan X., et al. (2023). Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 27351–27393. DOI: 10.52202/075280-1192 |
| [42] | Zhao W., Balachandran A., Tian C., et al. (2025). Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions. Belgrave D., Zhang C., Lin H., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)38: 4495–4530. DOI: 10.52202/085713-0159 |
| [43] | Zhang J., Zheng Y. and Qi D. (2017). Deep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction. Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 31:1655−1661. DOI:10.1609/aaai.v31i1.10735 |
| [44] | Shao Z., Zhang Z., Wang F., et al. (2022). Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Association for Computing Machinery) 1567−1577. DOI:10.1145/3534678.3539396 |
| [45] | Howard A., inversion, Makridakis S., et al. (2020). M5 Forecasting - Accuracy. |
| [46] | Cook A., DanB, inversion, et al. (2021). Store Sales - Time Series Forecasting. |
| [47] | Zhou J., Lu X., Xiao Y., et al. (2024). SDWPF: A Dataset for Spatial Dynamic Wind Power Forecasting over a Large Turbine Array. Sci. Data 11:649. DOI:10.1038/s41597-024-03427-5 |
| [48] | Kaltenborn J., Lange C., Ramesh V., et al. (2023). ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 21757–21792. DOI: 10.52202/075280-0952 |
| [49] | Chen W., Hao X., Wu Y., et al. (2024). Terra: A Multimodal Spatio-Temporal Dataset Spanning the Earth. Globerson A., Mackey L., Belgrave D., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)37: 66329–66356. DOI: 10.52202/079017-2121 |
| [50] | Ganem F., Vacaro L. B., Araujo E. C., et al. (2024). Mosqlimate: a platform to providing automatable access to data and forecasting models for arbovirus disease. DOI: 10.48550/arXiv.2410.18945 |
| [51] | Wahltinez O., Cheung A., Alcantara R., et al. (2022). COVID-19 Open-Data a global-scale spatially granular meta-dataset for coronavirus disease. Sci. Data 9:162. DOI:10.1038/s41597-022-01263-z |
| [52] | Zhang X., Ren G., Yu H., et al. (2025). LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence. DOI: 10.48550/arXiv.2509.03505 |
| [53] | Grinsztajn L., Flöge K., Key O., et al. (2025). TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models. DOI: 10.48550/arXiv.2511.08667 |
| [54] | Zhang X., Ren G., Yuan H., et al. (2026). LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence. arXiv:2609.17488. DOI:10.48550/arXiv.2609.17488 |
| [55] | Deng C., Yue Z. and Zhang Z. (2024). Polynormer: Polynomial-Expressive Graph Transformer in Linear Time. Kim B., Yue Y., Chaudhuri S., et al. (ed). International conference on learning representations 2024: 18323–18348. DOI: 10.48550/arXiv.2403.01232 |
| [56] | Rumelhart D. E., Hinton G. E. and Williams R. J. (1986). Learning representations by back-propagating errors. Nature 323:533−536. DOI:10.1038/323533a0 |
| [57] | Chen S. A., Li C. L., Yoder N. C., et al. (2023). TSMixer: An All-MLP Architecture for Time Series Forecasting. Trans. Mach. Learn. Res. DOI:10.48550/arXiv.2303.06053 |
| [58] | Cini A., Marisca I., Bianchi F. M., et al. (2023). Scalable Spatiotemporal Graph Neural Networks. Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 37:7218−7226. DOI:10.1609/aaai.v37i6.25880 |
| [59] | Goswami M., Szafer K., Choudhry A., et al. (2024). MOMENT: A Family of Open Time-series Foundation Models. Proceedings of the 41st International Conference on Machine Learning (PMLR) 235: 16115–16152. DOI:10.48550/arXiv.2402.03885 |
| Cheng K., Xu N., Zhang X., et al. (2026). When does heterogeneous structured evidence help? A unified benchmark for structured-data modeling. AI Plus 1:100015. https://doi.org/10.59717/ipj.aiplus.2026.100015 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
Example of entity-centric cross-modal prediction in a traffic sensor network
Normalized partial information decomposition (PID) across tasks