Обработка сканированных математических pdf-документов в сервисе семантического поиска формул

Main Article Content

Константин Сергеевич Николаев

Аннотация

Рассмотрена задача извлечения семантических связей между математическими формулами и концептами из неструктурированных PDF-документов, включая сканированные. Предложен сквозной конвейер обработки, объединяющий оптическое распознавание, восстановление структуры параграфов, выделение формул и переменных, а также их связывание с концептами на основе синтаксических шаблонов русского языка. Разработаны формальные модели для двух типов связей: «переменная – концепт» и «переменная – главная формула». Проведена экспериментальная оценка качества на коллекции из 150 математических статей; достигнуто повышение точности связывания на 18% по сравнению с предшествующей версией. Описаны архитектурные решения прототипа веб-сервиса, обеспечивающие асинхронную обработку и масштабируемость.

Article Details

Как цитировать
Николаев, К. С. «Обработка сканированных математических Pdf-документов в сервисе семантического поиска формул». Электронные библиотеки, т. 29, вып. 6, октябрь 2026 г., сс. 2555-72, doi:10.26907/1562-5419-2026-29-6-2555-2572.

Библиографические ссылки

1. Smith R. An overview of the tesseract OCR engine // Proc. Int. Conf. Doc. Anal. Recognition, ICDAR. 2007. Vol. 2. P. 629–633. https://doi.org/10.1109/ICDAR.2007.4376991
2. Constantin A., Pettifer S., Voronkov A. PDFX: Fully-automated PDF-to-XML conversion of scientific literature // DocEng 2013 – Proc. 2013 ACM Symp. Doc. Eng. 2013. P. 177–180. https://doi.org/10.1145/2494266.2494271
3. Ciancarini P., Iorio A. Di, Nuzzolese A. G., Peroni S., Vitali F. Semantic annotation of scholarly documents and citations // Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics). 2013. Vol. 8249 LNAI. P. 336–347. https://doi.org/10.1007/978-3-319-03524-6_29
4. Amethiya Y., Bajwa G. Automatic Table Detection and Tabular Data Extraction from Scanned Documents // Smart Innovation, Systems and Technologies. 2025. P. 75–87. https://doi.org/10.1007/978-981-96-0143-1_7
5. Bertin M., Atanassova I. Hybrid Approach for the Semantic Processing of Scientific Papers // Semant. Publ. Chall. Track 11 th Eur. Semant. Web Conf. (ESWC 2014). 2014. https://doi.org/10.1007/978-981-96-0143-1_7
6. Ahmad R., Afzal M. T., Qadir M. A. Information extraction from PDF sources based on rule-based system using integrated formats // Commun. Comput. Inf. Sci. 2016. Vol. 641. P. 293–308. https://doi.org/10.1007/978-3-319-46565-4_23
7. Schubotz M., Greiner-Petter A., Scharpf P., Meuschke N., Cohl H. S., Gipp B. Improving the Representation and Conversion of Mathematical Formulae by Considering their Textual Context // Proc. ACM/IEEE Jt. Conf. Digit. Libr. 2018. P. 233–242. https://doi.org/10.1145/3197026.3197058
8. Mathematical Markup Language (MathML) Version 3.0 2nd Edition. W3C Recommendation, 10 April 2014. URL: https://www.w3.org/TR/MathML3/
9. Greiner-Petter A., Youssef A., Ruas T., Miller B. R., Schubotz M., Aizawa A., Gipp B. Math-word embedding in math search and semantic extraction // Scientometrics. 2020. Vol. 125, No. 3. P. 3017–3046. https://doi.org/10.1007/s11192-020-03502-9
10. Nevzorova O., Kirillovich A., Nevzorov V., Nikolaev K. The semantic context models of mathematical formulas in scientific papers // CEUR Workshop Proc. 2018. Vol. 2277. P. 33–40.
11. Shah A. K., Dey A., Zanibbi R. A Math Formula Extraction and Evaluation Framework for PDF Documents // Lecture Notes in Computer Science. 2021. P. 19–34. https://doi.org/10.1007/978-3-030-86331-9_2
12. Siegel N., Lourie N., Power R., Ammar W. Extracting Scientific Figures with Distantly Supervised Neural Networks // Proc. ACM/IEEE Jt. Conf. Digit. Libr. 2018. P. 223–232. https://doi.org/10.1145/3197026.3197040
13. Semantic Scholar. URL: https://www.semanticscholar.org/
14. Neudecker C., Baierer K., Gerber M., Clausner C., Antonacopoulos A., Pletschacher S. A survey of OCR evaluation tools and metrics // The 6th International Workshop on Historical Document Imaging and Processing. New York, NY, USA: ACM, 2021. P. 13–18. https://doi.org/10.1145/3476887.3476888
15. Jaradeh M. Y., Oelen A., Farfar K. E., Prinz M., D'Souza J., Kismihók G., Stocker M., Auer S. Open research knowledge graph: Next generation infrastructure for semantic scholarly knowledge // K-CAP 2019 – Proceedings of the 10th International Conference on Knowledge Capture. New York, NY, USA: ACM, 2019. P. 243–246. https://doi.org/10.1145/3360901.3364435
16. Peroni S., Shotton D. FaBiO and CiTO: Ontologies for describing bibliographic resources and citations // J. Web Semant. 2012. Vol. 17. P. 33–43. https://doi.org/10.1016/j.websem.2012.08.001
17. Nikolaev K.S. Servis semanticheskogo poiska formul po kollekcii matematicheskih PDF-dokumentov // Informacionnye tehnologii i vychisli-tel'nye sistemy. 2025. № 3. S. 34–43. https://doi.org/10.14357/20718632250304
18. Kirillovich A. V., Nevzorova O. A., Lipachev E. K. OntoMathPRO 2.0 Ontology: Updates of Formal Model // Lobachevskii J. Math. 2022. Vol. 43, No. 12. P. 3504–3514. https://doi.org/10.1134/S1995080222150136
19. MSC2020 –Mathematics Subject Classification System. URL: https://msc2020.org/.
20. Celery –Distributed Task Queue. URL: https://docs.celeryq.dev/.
21. Redis – The real-time context engine for AI apps. URL: https://redis.io/.


Наиболее читаемые статьи этого автора (авторов)