# Convert PDF into TXT

**URL:** <https://forum.codewithmosh.com/t/convert-pdf-into-txt/19875>\
**Category:** Python\
**Created:** [April 26, 2023, 4:29pm UTC](https://forum.codewithmosh.com/t/convert-pdf-into-txt/19875 "2023-04-26T16:29:22Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![olivalej](https://avatars.discourse-cdn.com/v4/letter/o/ed655f/32.png) [@olivalej](https://forum.codewithmosh.com/u/olivalej)\
**Post date:** [April 26, 2023, 4:29pm UTC](https://forum.codewithmosh.com/t/convert-pdf-into-txt/19875/1 "2023-04-26T16:29:22Z")

</div>

Good day community,

I’m trying to compile some code to convert PDF to text, but the result is not what I expected. I have tried different libraries such as pytesseract, pdfminer, pdftotext, pdf2image, and OpenCV, but all of them extract the text incompletely or with errors. The last two codes that I used are these:

CODIGO 1  
import pytesseract  
from pdf2image import convert\_from\_path

# Configurar pytesseract

pytesseract.pytesseract.tesseract\_cmd = “/usr/bin/tesseract”  
pytesseract.pytesseract.tessdata\_dir\_config = ‘/usr/share/tesseract-ocr/4.00/tessdata’

# Ruta del archivo PDF

pdf\_path = “/content/drive/MyDrive/PDF/file.pdf” # Asegúrate de cambiar ‘tu\_archivo.pdf’ por el nombre real de tu archivo

# Convertir PDF a imágenes de alta calidad

images = convert\_from\_path(pdf\_path, dpi=300, fmt=“PNG”, thread\_count=4)

# Extraer texto de las imágenes

texts = [pytesseract.image\_to\_string(img, lang=“eng”, config=“–oem 1 --psm 11”) for img in images]

# Imprimir el texto extraído

for i, text in enumerate(texts):  
print(f"Texto de la página {i + 1}:\n{text}\n")

CODIGO 2  
from pdfminer.high\_level import extract\_text  
def convert\_pdf\_to\_txt(path):  
text = extract\_text(path)  
return text

# Cambia la ruta del archivo según la ubicación de tu archivo PDF

pdf\_path = ‘/content/drive/MyDrive/PDF/file.pdf’

# Convertir el PDF a texto

texto = convert\_pdf\_to\_txt(pdf\_path)

# Imprimir el texto en la consola

print(texto)

However, when I use online PDF to text converters, the conversion comes out very well, almost perfect, without the errors that I encounter in both codes. Here I attach the PDF that I want to convert to text and the results that I get from both codes when I try to convert my file.
