| 1 |
Zhong R, Chen Y H, Hu H, et al. Squirrel: testing database management systems with language validity and coverage feedback [C]//Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security. ACM, 2020: 955-970.
|
| 2 |
Li D, Liu Q, Guo Y, et al. BugForge: constructing and utilizing DBMS bug repository to enhance DBMS testing [PP/OL]. V1. arXiv (2026-04-03)[2026-07-29]. https://arxiv.org/abs/2604.03024.
|
| 3 |
Cui Z Y, Dou W S, Dai Q W, et al. Differentially testing database transactions for fun and profit [C]//Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. ACM, 2022: 1-12.
|
| 4 |
Dou W S, Cui Z Y, Dai Q W, et al. Detecting isolation bugs via transaction oracle construction [C]//Proceedings of the 2023 IEEE/ACM 45th International Conference on Software Engineering. IEEE, 2023: 1123-1135.
|
| 5 |
Rigger M, Su Z. Testing database engines via pivoted query synthesis [C]//Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation. 2020: 667-682.
|
| 6 |
Li K Q, Weng S Y, Liu P Y, et al. Leopard: a black-box approach for efficiently verifying various isolation levels [C]//Proceedings of the 2023 IEEE 39th International Conference on Data Engineering. IEEE, 2023: 722-735.
|
| 7 |
Tang X, Wu S, Zhang D, et al. Detecting logic bugs of join optimizations in DBMS [J]. Proceedings of the ACM on Management of Data, 2023, 1(1): 55.
|
| 8 |
Zhang Y L, Rodrigues K, Luo Y, et al. The inflection point hypothesis: a principled debugging approach for locating the root cause of a failure [C]//Proceedings of the 27th ACM Symposium on Operating Systems Principles. ACM, 2019: 131-146.
|
| 9 |
Ren X J, Wang S, Jin Z, et al. Relational debugging: pinpointing root causes of performance problems [C]//Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation. 2023: 65-80.
|
| 10 |
Liu D W, He C, Peng X, et al. MicroHECL: high-efficient root cause localization in large-scale microservice systems [C]//Proceedings of the 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice. IEEE, 2021: 338-347.
|
| 11 |
Zhou X, Li G, Sun Z, et al.. D-Bot: database diagnosis system using large language models. Proceedings of the VLDB Endowment, 2024, 17 (10): 2514- 2527.
|
| 12 |
Weng S Y, Wang Q S, Qu L Y, et al.. Lauca: a workload duplicator for benchmarking transactional database performance. IEEE Transactions on Knowledge and Data Engineering, 2024, 36 (7): 3180- 3194.
|
| 13 |
Binnig C, Kossmann D, Lo E, et al. QAGen: generating query-aware test databases [C]//Proceedings of the 2007 ACM SIGMOD International Conference on Management of Data. ACM, 2007: 341-352.
|
| 14 |
Wang Q S, Li H, Hu Z R, et al. Mirage: generating enormous databases for complex workloads [C]//Proceedings of the 2024 IEEE 40th International Conference on Data Engineering. IEEE, 2024: 3989-4001.
|
| 15 |
Huang S Y, Wang Z W, Zhang X Y, et al.. DBPA: a benchmark for transactional database performance anomalies. Proceedings of the ACM on Management of Data, 2023, 1 (1): 1- 26.
|
| 16 |
Colle R, Galanis L, Buranawatanachoke S, et al.. Oracle database replay. Proceedings of the VLDB Endowment, 2009, 2 (2): 1542- 1545.
|
| 17 |
Yu T, Zhang R, Yang K, et al. Spider: a large-scale human-labeled dataset for complex and cross-domain semantic parsing and Text-to-SQL task [C]//Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL, 2018: 3911-3921.
|
| 18 |
Li J, Hui B, Qu G, et al. Can LLM already serve as a database interface? A big bench for large-scale database grounded Text-to-SQLs [C]//Proceedings of the 37th Annual Conference on Neural Information Processing Systems. 2023, 36: 42330-42357.
|
| 19 |
Zhu Z J, Wang Y L, Liu D M. Hybrid query of Boolean filter and vector similarity search: benchmark, comparison and direction [C]//Proceedings of the 2025 5th Asia Conference on Information Engineering. IEEE, 2025: 117-124.
|
| 20 |
Shi J, Cai Y, Zheng W. Filtered approximate nearest neighbor search: a unified benchmark and systematic experimental study [Experiment, Analysis & Benchmark] [PP/OL]. V1. arXiv (2025-09-09)[2026-07-29]. https://arxiv.org/abs/2509.07789.
|
| 21 |
Ang E, Weldon S, Kim I K, et al. BranchBench: aligning database branching with agentic demands [PP/OL]. V1. arXiv (2026-04-19)[2026-07-29]. https://arxiv.org/abs/2604.17180.
|
| 22 |
Gao J L, Tao C Y, Bai J Y, et al. UniQL: towards dialect-universal benchmarking for Text-to-SQL [PP/OL]. V1. arXiv (2026-06-06)[2026-07-29]. https://arxiv.org/abs/2606.08018.
|
| 23 |
Zhou W, Li G, Wang H, et al. PARROT: a benchmark for evaluating LLMs in cross-system SQL translation [C]//Proceedings of the 39th Annual Conference on Neural Information Processing Systems. 2025: 38.
|
| 24 |
Pedro R, Coimbra M E, Castro D, et al. Prompt-to-SQL injections in LLM-integrated web applications: risks and defenses [C]//Proceedings of the 2025 IEEE/ACM International Conference on Software Engineering. 2025: 1768-1780.
|
| 25 |
Li D, Ling X, Li C, et al. AIA: autoregression-based injection attacks against Text2SQL models [C]//Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025: 18262-18270.
|
| 26 |
Zhan Q, Liang Z, Ying Z, et al. InjecAgent: benchmarking indirect prompt injections in tool-integrated large language model agents [C]//Proceedings of the Findings of the Association for Computational Linguistics. 2024: 10471-10506.
|