Logotipo del repositorio

Desarrollo de Fileoteca: aplicación web local y de código abierto para la gestión y clasificación automática de documentos digitales mediante técnicas bayesianas.

dc.contributor.advisorSeligmann Trujillo, Carlos David
dc.contributor.authorBARRERO OLIVEROS , JUAN SEBASTIAN
dc.coverage.temporal2026-1
dc.date.accessioned2026-07-08T14:10:45Z
dc.date.issued2026-07-03
dc.description.abstractEl proyecto Fileoteca consiste en el desarrollo de una aplicación web local y de código abierto orientada a la gestión, organización y clasificación de documentos digitales en múltiples formatos, como PDF, DOCX, XLSX, PPTX, imágenes y archivos de texto plano. La propuesta surge de la necesidad de contar con una herramienta que permita administrar documentos de manera estructurada desde el propio equipo del usuario, sin depender de servicios externos y manteniendo el control de la información. El sistema permite incorporar documentos mediante una interfaz web, la función de arrastrar y soltar, o desde el menú contextual del explorador de Windows. Una vez añadidos, los archivos pueden ser procesados para extraer su contenido textual mediante OCR o procesamiento directo, según el tipo de documento. Posteriormente, la aplicación facilita su organización mediante categorías, subcategorías, carpetas jerárquicas, etiquetas, favoritos y búsqueda de texto completo. Desde el punto de vista técnico, Fileoteca se desarrolló con una arquitectura modular compuesta por un backend en Golang con PocketBase, un servicio auxiliar en Python encargado del procesamiento documental y un frontend en Svelte para la interacción con el usuario. La comunicación entre componentes se realiza mediante gRPC. Además, el sistema incluye autenticación local basada en JWT, integración con el explorador de Windows y generación automática de miniaturas documentales. Conceptualmente, el proyecto integra principios de gestión documental digital, extracción de texto y clasificación probabilística mediante el algoritmo Naive Bayes. Este clasificador utiliza el contenido textual de los documentos para sugerir categorías dentro del sistema, incorporando un mecanismo de confianza que determina si la clasificación automática es adecuada o si requiere intervención del usuario. El desarrollo se realizó bajo un enfoque incremental, abarcando análisis, diseño, implementación, integración, validación y documentación. Finalmente, el sistema fue validado mediante 63 pruebas unitarias automatizadas, 12 pruebas funcionales manuales y una evaluación del clasificador, que obtuvo una exactitud global del 83,3 %. Como resultado, Fileoteca se consolida como una aplicación funcional, local y verificable.
dc.description.abstractenglishThe Fileoteca project involves the development of an open-source, local web application designed to manage, organize, and classify digital documents across various formats—such as PDF, DOCX, XLSX, PPTX, images, and plain text files. The initiative stems from the need for a tool capable of managing documents in a structured manner directly on the user's machine, eliminating reliance on external services while maintaining full control over the information. The system allows documents to be imported via a web interface, a drag-and-drop function, or the Windows Explorer context menu. Once added, files can be processed to extract their textual content using either OCR or direct processing, depending on the document type. The application then facilitates organization through categories, subcategories, hierarchical folders, tags, favorites, and full-text search capabilities. From a technical standpoint, Fileoteca utilizes a modular architecture comprising a Golang backend with PocketBase, a Python auxiliary service for document processing, and a Svelte frontend for user interaction. Component communication is handled via gRPC. Additionally, the system features local JWT-based authentication, Windows Explorer integration, and automatic document thumbnail generation. Conceptually, the project integrates principles of digital document management, text extraction, and probabilistic classification using the Naive Bayes algorithm. This classifier leverages the documents' textual content to suggest system categories, incorporating a confidence mechanism to determine whether the automatic classification is suitable or requires user intervention. Development followed an incremental approach, encompassing analysis, design, implementation, integration, validation, and documentation. Finally, the system was validated through 63 automated unit tests, 12 manual functional tests, and an evaluation of the classifier, which achieved an overall accuracy of 83.3%. Consequently, Fileoteca stands as a functional, local, and verifiable application.
dc.description.tableofcontents1. Resumen del proyecto 6 2. TÍTULO DE LA PROPUESTA 7 3. PLANTEAMIENTO DEL PROBLEMA 7 3.1. Descripción del problema 7 3.2. Elementos del problema 8 3.2.1. Documentos digitales no estructurados 8 3.2.2. Dependencia de organización manual 8 3.2.3. Ausencia de análisis de contenido 8 3.2.4. Limitaciones en la recuperación de información 8 3.2.5. Necesidad de soluciones locales 8 3.3. Formulación del problema 8 3.4. Consideraciones adicionales 9 4. OBJETIVOS 9 4.1. Objetivo general 9 4.2. Objetivos específicos 9 5. JUSTIFICACIÓN 10 6. MARCO TEÓRICO 11 6.1. Marco conceptual 11 6.1.1. Gestión documental digital 12 6.1.2. Clasificación de documentos basada en texto 13 6.1.3. Modelo Naive Bayes 15 6.1.4. Reconocimiento óptico de caracteres (OCR) 18 6.1.5. Arquitectura modular de software 19 6.1.6. Tecnologías utilizadas en el proyecto. 21 6.1.7. Modelo de datos del sistema 23 6.1.8. Seguridad y autenticación 25 6.2. Estado del arte 25 6.2.1. Sistemas comerciales de gestión documental 27 6.2.2. Sistemas de gestión documental de código abierto 30 6.2.3. Comparación de soluciones existentes 33 6.2.4. Técnicas de clasificación automática de documentos 36 7. METODOLOGÍA 39 7.1. Enfoque metodológico 39 7.2. Etapas del proyecto 40 7.2.1. Etapa 1: Análisis de requerimientos 40 7.2.2. Etapa 2: Diseño del sistema 40 7.2.3. Etapa 3: Implementación del sistema 40 7.2.4. Etapa 4: Pruebas 40 7.2.5. Etapa 5: Documentación. 40 7.3. Recolección y procesamiento de la información 40 7.4. Organización y análisis de los datos 44 7.5. Diseño del sistema como estrategia metodológica 44 7.6. Implementación del sistema 45 7.7. Integración de módulos. 46 7.8. Integración final del sistema 47 7.9. Validación funcional del sistema integrado 48 7.10. Evaluación del clasificador documental 50 8. RESULTADOS 51 8.1. Resultados obtenidos 51 8.2. Desarrollo e implementación del sistema 52 8.2.1. Implementación del backend 52 8.2.2. Implementación de la base de datos 52 8.2.3. Implementación del procesamiento documental 53 8.2.4. Implementación del mecanismo de clasificación 53 8.2.5. Desarrollo del frontend 53 8.2.6. Integración funcional del sistema 54 8.3. Validación y pruebas del sistema 54 8.3.1. Pruebas unitarias automatizadas 54 8.3.2. Pruebas funcionales integradas 55 8.3.3. Evaluación del clasificador 55 8.4. Cumplimiento de los objetivos específicos 55 8.5. Producto funcional desarrollado 56 9. CRONOGRAMA 57 9.1. Entregas académicas 59 9.2. Consideraciones del cronograma 60 9.3. Cumplimiento del cronograma 60 10. Conclusiones 60 11. Bibliografía 62 12. Anexos 65
dc.identifier.urihttps://hdl.handle.net/10823/8279
dc.relation.referencesAariff. (2025, noviembre 12). Classification of Text Documents using Naive Bayes. Medium. https://medium.com/%40lmohammedaariff/classification-of-text-documents-using-naive- bayes-9b67e1dcd3ef
dc.relation.referencesAggarwal, C. C., & Zhai, C. (2012). Mining text data. Springer.
dc.relation.referencesDevlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, 4171-4186.
dc.relation.referencesHotho, A., Nürnberger, A., & Paaß, G. (2005). A brief survey of text mining. LDV Forum, 20(1), 19-62.
dc.relation.referencesISO 15489-1:2016. (s. f.). ISO. https://www.iso.org/standard/62542.html
dc.relation.referencesJoachims, T. (1998). Text categorization with support vector machines: Learning with many relevant features. En Machine Learning: ECML-98 (pp. 137-142). Springer.
dc.relation.referencesJohn, R. (2025, marzo 27). Understanding document classification: A step-wise breakdown.
dc.relation.referencesKavlakoglu, E. (2025, noviembre 17). What are Naïve Bayes classifiers? Ibm.com. https://www.ibm.com/think/topics/naive-bayes
dc.relation.referencesManning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
dc.relation.referencesMcCallum, A., & Nigam, K. (1998). A comparison of event models for Naive Bayes text classification. Proceedings of the AAAI-98 Workshop on Learning for Text Categorization, 41-48.
dc.relation.referencesMinisterio de Comercio, Industria y Turismo. (2020). Programa de gestión de documentos electrónicos. https://www.mincit.gov.co/
dc.relation.referencesMori, S., Suen, C. Y., & Yamamoto, K. (1999). Historical review of OCR research and development. Proceedings of the IEEE, 80(7), 1029-1058.
dc.relation.referencesSebastiani, F. (2002). Machine learning in automated text categorization. ACM Computing Surveys, 34(1), 1–47. https://doi.org/10.1145/505282.505283
dc.relation.referencesSingh, H. (2020, enero 3). Naive Bayes: A Machine Learning Based Text Classifier.
dc.relation.referencesResearchgate.net. https://www.researchgate.net/publication/379512614_Naive_Bayes_A_Machine_Learning_Based_Text_Classifier
dc.relation.referencesSharma, N., Chen, M., & Sheth, B. (2006). Applied OCR techniques for document image processing. Proceedings of the International Conference on Computing: Techniques and Innovations.
dc.relation.referencesUniversitat Oberta de Catalunya. (2022). ¿Qué es un sistema de gestión documental (SGD)? https://www.uoc.edu/portal/es/arxiu/gestio-documental/que-es/index.html
dc.relation.referencesZhang, Z. (2016). Naïve Bayes classification in R. Annals of Translational Medicine, 4(12), 241. https://doi.org/10.21037/atm.2016.03.38
dc.relation.referencesAssociation for Intelligent Information Management. (2020). What is enterprise content management (ECM)? https://www.aiim.org
dc.relation.referencesDavenport, T. H. (2013). Process innovation: Reengineering work through information technology. Harvard Business Review Press.
dc.relation.referencesMicrosoft. (2025). SharePoint documentation. Microsoft Learn. https://learn.microsoft.com/sharepoint/
dc.relation.referencesM-Files Corporation. (2025). M-Files documentation. https://www.m-files.com
dc.relation.referencesDocuWare GmbH. (2025). DocuWare documentation. https://start.docuware.com
dc.relation.referencesOpenText Corporation. (2025). OpenText Document Management. https://www.opentext.com/products/document-management
dc.relation.referencesPaperless-ngx Contributors. (2025). Paperless-ngx documentation. https://docs.paperless-ngx.com
dc.relation.referencesMayan EDMS Contributors. (2025). Mayan EDMS documentation. https://docs.mayan-edms.com
dc.relation.referencesOpenKM. (2025). OpenKM documentation. https://docs.openkm.com
dc.relation.referencesSeedDMS Development Team. (2025). SeedDMS documentation. https://www.seeddms.org
dc.relation.referencesSmith, R. (2007). An overview of the Tesseract OCR engine. Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), 629–633. https://doi.org/10.1109/ICDAR.2007.4376991
dc.relation.referencesMori, S., Suen, C. Y., & Yamamoto, K. (1999). Historical review of OCR research and development. Proceedings of the IEEE, 80(7), 1029–1058. https://doi.org/10.1109/5.156468
dc.relation.referencesLewis, D. D. (1998). Naive (Bayes) at forty: The independence assumption in information retrieval. In Machine Learning: ECML-98 (pp. 4–15). Springer. https://doi.org/10.1007/BFb0026666
dc.relation.referencesMcCallum, A., & Nigam, K. (1998). A comparison of event models for Naive Bayes text classification. AAAI Workshop on Learning for Text Categorization. https://cdn.aaai.org/Workshops/1998/WS-98-05/WS98-05-007.pdf
dc.relation.referencesManning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press. https://nlp.stanford.edu/IR-book/
dc.relation.referencesCortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. https://doi.org/10.1007/BF00994018
dc.relation.referencesBreiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
dc.relation.referencesDevlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT. https://doi.org/10.48550/arXiv.1810.04805
dc.relation.referencesVaswani, A., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. https://doi.org/10.48550/arXiv.1706.03762
dc.relation.referencesJurafsky, D., & Martin, J. H. (2025). Speech and language processing (3rd ed., draft). Stanford University. https://web.stanford.edu/~jurafsky/slp3/
dc.subject.keywordsDocument Management
dc.subject.keywordsDocument Classification
dc.subject.keywordsNaive Bayes
dc.subject.keywordsOCR
dc.subject.keywordsLocal web application
dc.subject.keywordsGRPC
dc.subject.keywordsModular architecture
dc.subject.lembAplicaciones web
dc.subject.lembDesarrollo de sitios web
dc.subject.lembOrganización de la información
dc.subject.proposalGestión documental
dc.subject.proposalClasificación de documentos
dc.subject.proposalNaive Bayes
dc.subject.proposalOCR
dc.subject.proposalAplicación web local
dc.subject.proposalGRPC
dc.subject.proposalArquitectura modular
dc.titleDesarrollo de Fileoteca: aplicación web local y de código abierto para la gestión y clasificación automática de documentos digitales mediante técnicas bayesianas.
dc.title.translatedDevelopment of Fileoteca: A local, open-source web application for the management and automatic classification of digital documents using Bayesian techniques.
dc.typebachelorThesis
dc.type.coarhttp://purl.org/coar/resource_type/c_7a1f
dc.type.coarversionhttp://purl.org/coar/version/c_71e4c1898caa6e32
dc.type.driverinfo:eu-repo/semantics/bachelorThesis
dc.type.localTesis/Trabajo de grado - Monografía - Pregrado
dc.type.versioninfo:eu-repo/semantics/submittedVersion

Archivos

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
barrerooliverosjuansebastian_163091_95216899_TRABAJO DE GRADO - ENTREGABLE FINAL-2.docx
Tamaño:
5.26 MB
Formato:
Microsoft Word XML

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
license.txt
Tamaño:
1.71 KB
Formato:
Item-specific license agreed upon to submission
Descripción: