C# Extract text from PDF using PdfSharp

Took Sergio’s answer and made some extension methods. I also changed the accumulation of strings into an iterator. public static class PdfSharpExtensions { public static IEnumerable<string> ExtractText(this PdfPage page) { var content = ContentReader.ReadContent(page); var text = content.ExtractText(); return text; } public static IEnumerable<string> ExtractText(this CObject cObject) { if (cObject is COperator) { var cOperator … Read more

How to extract common / significant phrases from a series of text entries

I suspect you don’t just want the most common phrases, but rather you want the most interesting collocations. Otherwise, you could end up with an overrepresentation of phrases made up of common words and fewer interesting and informative phrases. To do this, you’ll essentially want to extract n-grams from your data and then find the … Read more

PDF Parsing Using Python – extracting formatted and plain texts [closed]

You can also take a look at PDFMiner (or for older versions of Python see PDFMiner and PDFMiner). A particular feature of interest in PDFMiner is that you can control how it regroups text parts when extracting them. You do this by specifying the space between lines, words, characters, etc. So, maybe by tweaking this … Read more

How to extract string following a pattern with grep, regex or perl [duplicate]

Since you need to match content without including it in the result (must match name=” but it’s not part of the desired result) some form of zero-width matching or group capturing is required. This can be done easily with the following tools: Perl With Perl you could use the n option to loop line by … Read more

How to extract a substring using regex

Assuming you want the part between single quotes, use this regular expression with a Matcher: “‘(.*?)'” Example: String mydata = “some string with ‘the data i want’ inside”; Pattern pattern = Pattern.compile(“‘(.*?)'”); Matcher matcher = pattern.matcher(mydata); if (matcher.find()) { System.out.println(matcher.group(1)); } Result: the data i want