KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I'm trying to use Python to processes some PDF forms that were filled out and signed using Adobe Acrobat Reader. I've tried: The pdfminer demo: it didn't dump any of the filled out data. pyPdf : it maxed a core for 2 minutes when I tried to load the file with PdfFileReader(f) and I just gave up and killed it. Jython and PDFBox : got that working great but the startup time is excessive, I'll just write an external utility in straight Java if that's my only option. I can keep hunting for libraries and trying them but I'm hoping someone already has an efficient solution for this. Update: Based on Steven's answer I looked into pdfminer and it did the trick nicely. from argparse import ArgumentParser import pickle import pprint from pdfminer.pdfparser import PDFParser from pdfminer.pdfdocument import PDFDocument from pdfminer.pdftypes import resolve1, PDFObjRef def load_form(filename): """Load pdf form contents into a nested list of name/value tuples""" with open(filename, 'rb') as file: parser = PDFParser(file) doc = PDFDocument(parser) return [load_fields(resolve1(f)) for f in resolve1(doc.catalog['AcroForm'])['Fields']] def load_fields(field): """Recursively load form fields""" form = field.get('Kids', None) if form: return [load_fields(resolve1(f)) for f in form] else: # Some field types, like signatures, need extra resolving return (field.get('T').decode('utf-16'), resolve1(field.get('V'))) def parse_cli(): """Load command line arguments""" parser = ArgumentParse
Tags (comma-separated)
Save Edits
Cancel