| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
tabula-sharp is a library for extracting tables from PDF files — it is a port of tabula-java
NuGet packages available on the releases page and on www.nuget.org:
using (PdfDocument document = PdfDocument.Open("doc.pdf", new ParsingOptions() { ClipPaths = true }))
{
PageArea page = ObjectExtractor.Extract(document, 1);
// detect canditate table zones
SimpleNurminenDetectionAlgorithm detector = new SimpleNurminenDetectionAlgorithm();
var regions = detector.Detect(page);
IExtractionAlgorithm ea = new BasicExtractionAlgorithm();
IReadOnlyList<Table> tables = ea.Extract(page.GetArea(regions[0].BoundingBox)); // take first candidate area
var table = tables[0];
var rows = table.Rows;
}using (PdfDocument document = PdfDocument.Open("doc.pdf", new ParsingOptions() { ClipPaths = true }))
{
PageArea page = ObjectExtractor.Extract(document, 1);
IExtractionAlgorithm ea = new SpreadsheetExtractionAlgorithm();
IReadOnlyList<Table> tables = ea.Extract(page);
var table = tables[0];
var rows = table.Rows;
}| Back | FazBrowse Home | New Git URL |