Automated analysis of articles and scientific and technical reports is complicated by the content of formulas, graphs, images, and tables. When using text analysis methods, the information contained in these components is completely or partially lost. Artificial intelligence systems can sometimes analyze and formulate in text what is depicted on the graph and image. However, this approach requires computational resources, and its accuracy and reliability may be low, especially if it is not a known graph, but a new result. And there is no way to verify the result. As part of our system, text analysis is performed for subsequent processing by a human operator. But when processing text, graphs, images, and formulas are lost, and tables lose their structure, as the text undergoes multi-stage processing, word breakdown, and conversion to its initial form. Additional modules are needed to extract, store and process such elements, allowing them to maintain their connection with the document at all stages of work. A promising approach is to change the original way of creating and forming scientific documents, using many years of experience in creating web resources, taking into account their automated processing.