Splitting string into multiple rows in Oracle

This may be an improved way (also with regexp and connect by): with temp as ( select 108 Name, ‘test’ Project, ‘Err1, Err2, Err3’ Error from dual union all select 109, ‘test2’, ‘Err1’ from dual ) select distinct t.name, t.project, trim(regexp_substr(t.error, ‘[^,]+’, 1, levels.column_value)) as error from temp t, table(cast(multiset(select level from dual connect by … Read more

Can a line of Python code know its indentation nesting level?

If you want indentation in terms of nesting level rather than spaces and tabs, things get tricky. For example, in the following code: if True: print( get_nesting_level()) the call to get_nesting_level is actually nested one level deep, despite the fact that there is no leading whitespace on the line of the get_nesting_level call. Meanwhile, in … Read more

How to get rid of punctuation using NLTK tokenizer?

Take a look at the other tokenizing options that nltk provides here. For example, you can define a tokenizer that picks out sequences of alphanumeric characters as tokens and drops everything else: from nltk.tokenize import RegexpTokenizer tokenizer = RegexpTokenizer(r’\w+’) tokenizer.tokenize(‘Eighty-seven miles to go, yet. Onward!’) Output: [‘Eighty’, ‘seven’, ‘miles’, ‘to’, ‘go’, ‘yet’, ‘Onward’]

Scanner vs. StringTokenizer vs. String.Split

They’re essentially horses for courses. Scanner is designed for cases where you need to parse a string, pulling out data of different types. It’s very flexible, but arguably doesn’t give you the simplest API for simply getting an array of strings delimited by a particular expression. String.split() and Pattern.split() give you an easy syntax for … Read more

Looking for a clear definition of what a “tokenizer”, “parser” and “lexers” are and how they are related to each other and used?

A tokenizer breaks a stream of text into tokens, usually by looking for whitespace (tabs, spaces, new lines). A lexer is basically a tokenizer, but it usually attaches extra context to the tokens — this token is a number, that token is a string literal, this other token is an equality operator. A parser takes … Read more

What is the easiest/best/most correct way to iterate through the characters of a string in Java?

I use a for loop to iterate the string and use charAt() to get each character to examine it. Since the String is implemented with an array, the charAt() method is a constant time operation. String s = “…stuff…”; for (int i = 0; i < s.length(); i++){ char c = s.charAt(i); //Process char } … Read more

How do I tokenize a string in C++?

The Boost tokenizer class can make this sort of thing quite simple: #include <iostream> #include <string> #include <boost/foreach.hpp> #include <boost/tokenizer.hpp> using namespace std; using namespace boost; int main(int, char**) { string text = “token, test string”; char_separator<char> sep(“, “); tokenizer< char_separator<char> > tokens(text, sep); BOOST_FOREACH (const string& t, tokens) { cout << t << “.” … Read more