10 ms·
Looking at the code, it looks like they used existing Python packages to read and parse MS Office formats, not what I expected, seeing that the repo is in Micro
by wis 2y ago
Looking at the code, it looks like they used existing Python packages to read and parse MS Office formats, not what I expected, seeing that the repo is in Microsoft's org on GitHub I expected them to have used Microsoft's "official" libraries for parsing these formats, through Component Object Model (COM).
They used Mammoth for docx (Word) [1][2]
Python-pptx for ppt (PowerPoint) [3][4]
and Pandas for XSLX (Excel) [5]
[1] https://github.com/microsoft/markitdown/blob/70ab149ff1657c327ebd6ca940988f5d3c5d80d0/src/markitdown/_markitdown.py#L495 https://github.com/microsoft/markitdown/blob/70ab149ff1657c3... [2] https://pypi.org/project/mammoth/ https://pypi.org/project/mammoth/
[3] https://github.com/microsoft/markitdown/blob/70ab149ff1657c327ebd6ca940988f5d3c5d80d0/src/markitdown/_markitdown.py#L539 https://github.com/microsoft/markitdown/blob/70ab149ff1657c3...
[4] https://pypi.org/project/python-pptx/ https://pypi.org/project/python-pptx/
[5] https://github.com/microsoft/markitdown/blob/70ab149ff1657c327ebd6ca940988f5d3c5d80d0/src/markitdown/_markitdown.py#L513 https://github.com/microsoft/markitdown/blob/70ab149ff1657c3...
- jamwil 2y agoCOM requires you to interact with the files through the associated MS Office applications, whereas these libs parse the ooxml file format directly.