| 1 |
Radford A, Narasimhan K, Salimans T, et al. Improving language understanding by generative pre-training [R]. OpenAI, 2018.
|
| 2 |
Radford A, Wu J, Child R, et al. Language models are unsupervised multitask learners [R]. OpenAI, 2019.
|
| 3 |
Brown T B, Mann B, Ryder N, et al. Language models are few-shot learners [C]//Proceedings of the 34th International Conference on Neural Information Processing Systems. ACM, 2020: 1877-1901.
|
| 4 |
Zhang K C, Li J, Li G, et al. CodeAgent: enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges [C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2024: 13643-13658.
|
| 5 |
Abramson J, Adler J, Dunger J, et al.. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 2024, 630, 493- 500.
|
| 6 |
McDuff D, Schaekermann M, Tu T, et al.. Towards accurate differential diagnosis with large language models. Nature, 2025, 642 (8067): 451- 457.
|
| 7 |
Katz D M, Bommarito M J, Gao S, et al.. GPT-4 passes the bar exam. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 2024, 382 (2270): 20230254.
|
| 8 |
Bengio Y, Hinton G, Yao A, et al.. Managing extreme AI risks amid rapid progress. Science, 2024, 384 (6698): 842- 845.
|
| 9 |
Karpathy A. 2025 LLM Year in Review [EB/OL]. (2025-12-20)[2026-07-10]. https://karpathy.bearblog.dev/year-in-review-2025/
|
| 10 |
Kwon W, Li Z H, Zhuang S Y, et al. Efficient memory management for large language model serving with PagedAttention [C]//Proceedings of the 29th Symposium on Operating Systems Principles. ACM, 2023: 611-626.
|
| 11 |
Zheng L M, Yin L S, Xie Z Q, et al. SGLang: efficient execution of structured language model programs [C]//Advances in Neural Information Processing Systems 37. 2024: 62557-62583.
|
| 12 |
Krizhevsky A, Sutskever I, Hinton G E.. ImageNet classification with deep convolutional neural networks. Communications of the ACM, 2017, 60 (6): 84- 90.
|
| 13 |
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need [C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. ACM, 2017: 6000-6010.
|
| 14 |
Sutton R. The bitter lesson [EB/OL]. (2019-03-13) [2026-07-10]. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
|
| 15 |
DeepSeek-AI, Liu A X, Feng B, et al. DeepSeek-V3 technical report [PP/OL]. V2. arXiv (2025-02-18) [2026-07-10]. https://doi.org/10.48550/arXiv.2412.19437.
|
| 16 |
Zhou C T, Liu P F, Xu P X, et al. LIMA: less is more for alignment [C]//Advances in Neural Information Processing Systems 36 (NeurIPS 2023). 2023: 55006-55021.
|
| 17 |
Lin B Y, Ravichander A, Lu X M, et al. The unlocking spell on base LLMs: rethinking alignment via in-context learning [C]//International Conference on Learning Representations. 2024: 24907-24933.
|
| 18 |
Liu Z C, Chen C, Li W, et al. Understanding R1-zero-like training: a critical perspective [C]//Proceedings of the Second Conference on Language Modeling. 2025. DOI:10.48550/arXiv.2503.20783.
|
| 19 |
Yue Y, Chen Z Q, Lu R, et al. Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? [C]//Advances in Neural Information Processing Systems 38 Main Conference (NeurIPS 2025). 2025: 57654-57689.
|
| 20 |
Hendrycks D, Song D, Szegedy C, et al. A definition of AGI [PP/OL]. V3. arXiv (2025-12-03) [2026-07-10]. https://doi.org/10.48550/arXiv.2510.18212.
|
| 21 |
Li J T, Hu R J, Huang K Z, et al. PertEval: unveiling real knowledge capacity of LLMs with knowledge-invariant perturbations [C]//Advances in Neural Information Processing Systems 37. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024: 10679-10706.
|
| 22 |
Xu Y Y, Hu R J, Ying H C, et al. Large language models could be rote learners [PP/OL]. V5. arXiv (2026-05-15) [2026-07-10]. https://doi.org/10.48550/arXiv.2504.08300.
|
| 23 |
Borisov V, Leemann T, Seßler K, et al.. Deep neural networks and tabular data: a survey. IEEE Transactions on Neural Networks and Learning Systems, 2024, 35 (6): 7499- 7519.
|
| 24 |
Chen T Q, Guestrin C. XGBoost: a scalable tree boosting system [C]//Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 2016: 785-794.
|
| 25 |
Cheng Y, Hu R J, Ying H C, et al. Arithmetic feature interaction is necessary for deep tabular learning [C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(10): 11516-11524.
|
| 26 |
Yuan Y F, Li J T, Zhang W J, et al. Summarize-exemplify-reflect: data-driven insight distillation empowers LLMs for few-shot tabular classification [C]//Findings of the Association for Computational Linguistics: EMNLP 2025. Association for Computational Linguistics. 2025: 12324-12348.
|
| 27 |
Zhang Y B, Ye H H, Bai Y, et al. Automating complex document workflows via stepwise and rollback-enabled operation orchestration [C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(43): 36518-36526.
|
| 28 |
Pancha N, Zhai A, Leskovec J, et al. PinnerFormer: sequence modeling for user representation at Pinterest [C]//Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ACM, 2022: 3702-3712.
|
| 29 |
Botta E, Yang J, Hsu Y P, et al. PinRec: unified generative retrieval for Pinterest recommender systems [PP/OL]. V7. arXiv (2026-06-18) [2026-07-10]. https://doi.org/10.48550/arXiv.2504.10507.
|