Article Contents
ARTICLE   Open Access     Cite

When does heterogeneous structured evidence help? A unified benchmark for structured-data modeling

    Show all affliationsShow less
More Information
  • DownLoad: Full size image
    1. Introduce an entity-centric Trimodal Structured-Data Benchmark (TSDBench) for multimodal learning.

      TSDBench validates multimodal information gains via ablation and partial information decomposition (PID).

      Results reveal cross-modal gains are task-dependent, asymmetric, and vary with prediction targets.

      Flow-PID exposes gaps between theoretical multimodal synergy and what current models can exploit.

  • Real-world entities are often characterized by heterogeneous structured information, including static attributes, temporal dynamics, and relational dependencies. However, existing benchmarks provide limited understanding of whether multimodal models can effectively exploit complementary information across these modalities. In this work, we introduce Trimodal Structured-Data Benchmark (TSDBench), an entity-centric benchmark for evaluating multimodal structured-data learning. To establish a reliable foundation for evaluation, we first verify the multimodal complementarity of benchmark tasks from both empirical and theoretical perspectives: input-ablation experiments demonstrate widespread predictive gains from incorporating additional modalities, while Flow-Partial Information Decomposition (Flow-PID) analysis quantifies the unique, redundant, and synergistic information contributed by different modalities. Furthermore, TSDBench adopts a target-rotation design that enables systematic analysis of cross-modal interactions by allowing different static features or time series signals to alternately serve as prediction targets. Extensive experiments reveal that multimodal gains are highly task-dependent and asymmetric, varying with the target modality and dataset characteristics. Although the proposed fusion models achieve strong overall performance, Flow-PID analysis exposes a substantial gap between the complementary information available in multimodal data and what current architectures can effectively exploit. These findings highlight the need for more adaptive multimodal learning methods capable of selectively leveraging heterogeneous structured evidence.
  • 加载中
  • [1] Dua D. and Graff C. (2019). UCI Machine Learning Repository.

    View in Article Google Scholar

    [2] Vanschoren J., van Rijn J. N., Bischl B., et al. (2013). OpenML: networked science in machine learning. ACM SIGKDD Explor. Newsl. 15:49−60. DOI:10.1145/2641190.2641198

    View in Article CrossRef Google Scholar

    [3] Chen T. and Guestrin C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Association for Computing Machinery) 785−794. DOI:10.1145/2939672.2939785

    View in Article CrossRef Google Scholar

    [4] Prokhorenkova L., Gusev G., Vorobev A., et al. (2018). CatBoost: unbiased boosting with categorical features. Bengio S., Wallach H., Larochelle H., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)31: 6638–6648. DOI: 10.48550/arXiv.1706.09516

    View in Article Google Scholar

    [5] Gorishniy Y., Rubachev I., Khrulkov V., et al. (2021). Revisiting Deep Learning Models for Tabular Data. Ranzato M., Beygelzimer A., Dauphin Y., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)34: 18932–18943. DOI: 10.48550/arXiv.2106.11959

    View in Article Google Scholar

    [6] Hollmann N., Müller S., Eggensperger K., et al. (2023). TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second. International conference on learning representations. DOI:10.48550/arXiv.2207.01848

    View in Article CrossRef Google Scholar

    [7] Grinsztajn L., Oyallon E. and Varoquaux G. (2022). Why do tree-based models still outperform deep learning on typical tabular data? Koyejo S., Mohamed S., Agarwal A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)35: 507–520. DOI: 10.52202/068431-0037

    View in Article Google Scholar

    [8] Makridakis S., Spiliotis E. and Assimakopoulos V. (2020). The M4 Competition: 100, 000 time series and 61 forecasting methods. Int. J. Forecast. 36:54−74. DOI:10.1016/j.ijforecast.2019.04.014

    View in Article CrossRef Google Scholar

    [9] Godahewa R., Bergmeir C., Webb G. I., et al. (2021). Monash Time Series Forecasting Archive. Vanschoren J. and Yeung S. (ed). Proceedings of the neural information processing systems track on datasets and benchmarks 1. DOI: 10.48550/arXiv.2105.06643

    View in Article Google Scholar

    [10] Zhou H., Zhang S., Peng J., et al. (2021). Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 35:11106−11115. DOI:10.1609/aaai.v35i12.17325

    View in Article CrossRef Google Scholar

    [11] Wu H., Xu J., Wang J., et al. (2021). Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. Ranzato M., Beygelzimer A., Dauphin Y., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)34: 22419–22430. DOI: 10.48550/arXiv.2106.13008

    View in Article Google Scholar

    [12] Zeng A., Chen M., Zhang L., et al. (2023). Are Transformers Effective for Time Series Forecasting? Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 37:11121−11128. DOI:10.1609/aaai.v37i9.26317

    View in Article CrossRef Google Scholar

    [13] Nie Y., Nguyen N. H., Sinthong P., et al. (2023). A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. International conference on learning representations. DOI:10.48550/arXiv.2211.14730

    View in Article CrossRef Google Scholar

    [14] Liu Y., Hu T., Zhang H., et al. (2024). iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. Kim B., Yue Y., Chaudhuri S., et al. (ed). International conference on learning representations 2024: 11116–11140. DOI: 10.48550/arXiv.2310.06625

    View in Article Google Scholar

    [15] Kipf T. N. and Welling M. (2017). Semi-Supervised Classification with Graph Convolutional Networks. International conference on learning representations. DOI:10.48550/arXiv.1609.02907

    View in Article CrossRef Google Scholar

    [16] Hamilton W. L., Ying R. and Leskovec J. (2017). Inductive Representation Learning on Large Graphs. Guyon I., Luxburg U. V., Bengio S., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)30: 1024–1034. DOI: 10.48550/arXiv.1706.02216

    View in Article Google Scholar

    [17] Veličković P., Cucurull G., Casanova A., et al. (2018). Graph Attention Networks. International conference on learning representations. DOI:10.48550/arXiv.1710.10903

    View in Article CrossRef Google Scholar

    [18] Xu K., Hu W., Leskovec J., et al. (2019). How Powerful are Graph Neural Networks? International conference on learning representations. DOI:10.48550/arXiv.1810.00826

    View in Article CrossRef Google Scholar

    [19] Hu W., Fey M., Zitnik M., et al. (2020). Open Graph Benchmark: Datasets for Machine Learning on Graphs. Larochelle H., Ranzato M., Hadsell R., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)33: 22118–22133. DOI: 10.48550/arXiv.2005.00687

    View in Article Google Scholar

    [20] Huang S., Poursafaei F., Danovitch J., et al. (2023). Temporal Graph Benchmark for Machine Learning on Temporal Graphs. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 2056–2073. DOI: 10.52202/075280-0099

    View in Article Google Scholar

    [21] Bazhenov G., Platonov O. and Prokhorenkova L. (2025). GraphLand: Evaluating Graph Machine Learning Models on Diverse Industrial Data. Belgrave D., Zhang C., Lin H., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)38: 110211–110237. DOI: 10.52202/085713-3321

    View in Article Google Scholar

    [22] Fu Y., Shao Z., Yu C., et al. (2026). Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis. DOI: 10.48550/arXiv.2607.01918

    View in Article Google Scholar

    [23] Yu C., Wang F., Shao Z., et al. (2025). GinAR+: A Robust End-to-End Framework for Multivariate Time Series Forecasting With Missing Values. IEEE Trans. Knowl. Data Eng. 37:4635−4648. DOI:10.1109/tkde.2025.3569649

    View in Article CrossRef Google Scholar

    [24] Shao Z., Wang F., Xu Y., et al. (2025). Exploring Progress in Multivariate Time Series Forecasting: Comprehensive Benchmarking and Heterogeneity Analysis. IEEE Trans. Knowl. Data Eng. 37:291−305. DOI:10.1109/tkde.2024.3484454

    View in Article CrossRef Google Scholar

    [25] Li Y., Yu R., Shahabi C., et al. (2018). Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. International conference on learning representations. DOI:10.48550/arXiv.1707.01926

    View in Article CrossRef Google Scholar

    [26] Yu B., Yin H. and Zhu Z. (2018). Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. Proceedings of the twenty-seventh international joint conference on artificial intelligence (International Joint Conferences on Artificial Intelligence Organization) 3634−3640. DOI:10.24963/ijcai.2018/505

    View in Article CrossRef Google Scholar

    [27] Wu Z., Pan S., Long G., et al. (2019). Graph WaveNet for Deep Spatial-Temporal Graph Modeling. Proceedings of the twenty-eighth international joint conference on artificial intelligence (International Joint Conferences on Artificial Intelligence Organization) 1907−1913. DOI:10.24963/ijcai.2019/264

    View in Article CrossRef Google Scholar

    [28] Bai L., Yao L., Li C., et al. (2020). Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting. Larochelle H., Ranzato M., Hadsell R., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)33: 17804–17815. DOI: 10.48550/arXiv.2007.02842

    View in Article Google Scholar

    [29] Liu X., Xia Y., Liang Y., et al. (2023). LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 75354–75371. DOI: 10.52202/075280-3293

    View in Article Google Scholar

    [30] Fey M., Hu W., Huang K., et al. (2024). Position: Relational Deep Learning - Graph Representation Learning on Relational Databases. Salakhutdinov R., Kolter Z., Heller K., et al. (ed). Proceedings of the 41st International Conference on Machine Learning (PMLR)235: 13592–13607. DOI: 10.48550/arXiv.2312.04615

    View in Article Google Scholar

    [31] Robinson J., Ranjan R., Hu W., et al. (2024). RelBench: A Benchmark for Deep Learning on Relational Databases. Globerson A., Mackey L., Belgrave D., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)37: 21330–21341. DOI: 10.52202/079017-0672

    View in Article Google Scholar

    [32] Hu W., Yuan Y., Zhang Z., et al. (2024). PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning. DOI: 10.48550/arXiv.2404.00776

    View in Article Google Scholar

    [33] Baltrušaitis T., Ahuja C. and Morency L. P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41:423−443. DOI:10.1109/tpami.2018.2798607

    View in Article CrossRef Google Scholar

    [34] Zadeh A., Chen M., Poria S., et al. (2017). Tensor Fusion Network for Multimodal Sentiment Analysis. Proceedings of the 2017 conference on empirical methods in natural language processing (Association for Computational Linguistics) 1103−1114. DOI:10.18653/v1/d17-1115

    View in Article CrossRef Google Scholar

    [35] Liu Z., Shen Y., Lakshminarasimhan V. B., et al. (2018). Efficient Low-rank Multimodal Fusion With Modality-Specific Factors. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Association for Computational Linguistics) 2247−2256. DOI:10.18653/v1/p18-1209

    View in Article CrossRef Google Scholar

    [36] Tsai Y. H. H., Bai S., Liang P. P., et al. (2019). Multimodal Transformer for Unaligned Multimodal Language Sequences. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Association for Computational Linguistics) 6558−6569. DOI:10.18653/v1/p19-1656

    View in Article CrossRef Google Scholar

    [37] Jaegle A., Gimeno F., Brock A., et al. (2021). Perceiver: General Perception with Iterative Attention. Meila M. and Zhang T. (ed). Proceedings of the 38th International Conference on Machine Learning (PMLR)139: 4651–4664. DOI: 10.48550/arXiv.2103.03206

    View in Article Google Scholar

    [38] Liang P. P., Lyu Y., Fan X., et al. (2021). MultiBench: Multiscale Benchmarks for Multimodal Representation Learning. Vanschoren J. and Yeung S. (ed). Proceedings of the neural information processing systems track on datasets and benchmarks 1. DOI: 10.48550/arXiv.2107.07502

    View in Article Google Scholar

    [39] Williams P. L. and Beer R. D. (2010). Nonnegative Decomposition of Multivariate Information. DOI: 10.48550/arXiv.1004.2515

    View in Article Google Scholar

    [40] Bertschinger N., Rauh J., Olbrich E., et al. (2014). Quantifying Unique Information. Entropy 16:2161−2183. DOI:10.3390/e16042161

    View in Article CrossRef Google Scholar

    [41] Liang P. P., Cheng Y., Fan X., et al. (2023). Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 27351–27393. DOI: 10.52202/075280-1192

    View in Article Google Scholar

    [42] Zhao W., Balachandran A., Tian C., et al. (2025). Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions. Belgrave D., Zhang C., Lin H., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)38: 4495–4530. DOI: 10.52202/085713-0159

    View in Article Google Scholar

    [43] Zhang J., Zheng Y. and Qi D. (2017). Deep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction. Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 31:1655−1661. DOI:10.1609/aaai.v31i1.10735

    View in Article CrossRef Google Scholar

    [44] Shao Z., Zhang Z., Wang F., et al. (2022). Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Association for Computing Machinery) 1567−1577. DOI:10.1145/3534678.3539396

    View in Article CrossRef Google Scholar

    [45] Howard A., inversion, Makridakis S., et al. (2020). M5 Forecasting - Accuracy.

    View in Article Google Scholar

    [46] Cook A., DanB, inversion, et al. (2021). Store Sales - Time Series Forecasting.

    View in Article Google Scholar

    [47] Zhou J., Lu X., Xiao Y., et al. (2024). SDWPF: A Dataset for Spatial Dynamic Wind Power Forecasting over a Large Turbine Array. Sci. Data 11:649. DOI:10.1038/s41597-024-03427-5

    View in Article CrossRef Google Scholar

    [48] Kaltenborn J., Lange C., Ramesh V., et al. (2023). ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning. Oh A., Naumann T., Globerson A., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)36: 21757–21792. DOI: 10.52202/075280-0952

    View in Article Google Scholar

    [49] Chen W., Hao X., Wu Y., et al. (2024). Terra: A Multimodal Spatio-Temporal Dataset Spanning the Earth. Globerson A., Mackey L., Belgrave D., et al. (ed). Advances in neural information processing systems (Curran Associates, Inc.)37: 66329–66356. DOI: 10.52202/079017-2121

    View in Article Google Scholar

    [50] Ganem F., Vacaro L. B., Araujo E. C., et al. (2024). Mosqlimate: a platform to providing automatable access to data and forecasting models for arbovirus disease. DOI: 10.48550/arXiv.2410.18945

    View in Article Google Scholar

    [51] Wahltinez O., Cheung A., Alcantara R., et al. (2022). COVID-19 Open-Data a global-scale spatially granular meta-dataset for coronavirus disease. Sci. Data 9:162. DOI:10.1038/s41597-022-01263-z

    View in Article CrossRef Google Scholar

    [52] Zhang X., Ren G., Yu H., et al. (2025). LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence. DOI: 10.48550/arXiv.2509.03505

    View in Article Google Scholar

    [53] Grinsztajn L., Flöge K., Key O., et al. (2025). TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models. DOI: 10.48550/arXiv.2511.08667

    View in Article Google Scholar

    [54] Zhang X., Ren G., Yuan H., et al. (2026). LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence. arXiv:2609.17488. DOI:10.48550/arXiv.2609.17488

    View in Article Google Scholar

    [55] Deng C., Yue Z. and Zhang Z. (2024). Polynormer: Polynomial-Expressive Graph Transformer in Linear Time. Kim B., Yue Y., Chaudhuri S., et al. (ed). International conference on learning representations 2024: 18323–18348. DOI: 10.48550/arXiv.2403.01232

    View in Article Google Scholar

    [56] Rumelhart D. E., Hinton G. E. and Williams R. J. (1986). Learning representations by back-propagating errors. Nature 323:533−536. DOI:10.1038/323533a0

    View in Article CrossRef Google Scholar

    [57] Chen S. A., Li C. L., Yoder N. C., et al. (2023). TSMixer: An All-MLP Architecture for Time Series Forecasting. Trans. Mach. Learn. Res. DOI:10.48550/arXiv.2303.06053

    View in Article Google Scholar

    [58] Cini A., Marisca I., Bianchi F. M., et al. (2023). Scalable Spatiotemporal Graph Neural Networks. Proceedings of the AAAI conference on artificial intelligence (Association for the Advancement of Artificial Intelligence (AAAI)) 37:7218−7226. DOI:10.1609/aaai.v37i6.25880

    View in Article CrossRef Google Scholar

    [59] Goswami M., Szafer K., Choudhry A., et al. (2024). MOMENT: A Family of Open Time-series Foundation Models. Proceedings of the 41st International Conference on Machine Learning (PMLR) 235: 16115–16152. DOI:10.48550/arXiv.2402.03885

    View in Article CrossRef Google Scholar

  • Cite this article:

    Cheng K., Xu N., Zhang X., et al. (2026). When does heterogeneous structured evidence help? A unified benchmark for structured-data modeling. AI Plus 1:100015. https://doi.org/10.59717/ipj.aiplus.2026.100015
    Cheng K., Xu N., Zhang X., et al. (2026). When does heterogeneous structured evidence help? A unified benchmark for structured-data modeling. AI Plus 1:100015. https://doi.org/10.59717/ipj.aiplus.2026.100015

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(2)     Tables(6)

Supplementary Information

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(48) PDF downloads(15)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint