| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Forked from jsfenfen/whatwordwhere
Tooling to extract data from scanned paper forms OCR-ed by Tesseract using the HOCR standard.
experimenting with pdf2text and python pdf-table-extract
This project will liberate data from pdf files found on http://www.cityofjerseycity.com/pub-info.aspx?id=2430 and will create .csv and .json files to be uploaded on https://data.openjerseycity.org/…
Data from the United States Agency for International Development (USAID) Development Experience Clearinghouse (DEC).
Tooling to extract data from scanned paper forms OCR-ed by Tesseract using the HOCR standard.
Tools for working with Optical Character Recognition output
This project will liberate data from pdf files found on http://www.cityofjerseycity.com/pub-info.aspx?id=2430 and will create .csv and .json files to be uploaded on https://data.openjerseycity.org/dataset/jersey-city-2013-budget-adopted-spending
This uses regular expressions (in php, but can be any language) get data from the NYC EDC newsletters
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |