public final class WikipediaTokenizer extends Tokenizer
AttributeSource.State| Modifier and Type | Field and Description |
|---|---|
static int |
ACRONYM_ID |
static int |
ALPHANUM_ID |
static int |
APOSTROPHE_ID |
static String |
BOLD |
static int |
BOLD_ID |
static String |
BOLD_ITALICS |
static int |
BOLD_ITALICS_ID |
static int |
BOTH
Output the both the untokenized token and the splits
|
static String |
CATEGORY |
static int |
CATEGORY_ID |
static String |
CITATION |
static int |
CITATION_ID |
static int |
CJ_ID |
static int |
COMPANY_ID |
static int |
EMAIL_ID |
static String |
EXTERNAL_LINK |
static int |
EXTERNAL_LINK_ID |
static String |
EXTERNAL_LINK_URL |
static int |
EXTERNAL_LINK_URL_ID |
static String |
HEADING |
static int |
HEADING_ID |
static int |
HOST_ID |
static String |
INTERNAL_LINK |
static int |
INTERNAL_LINK_ID |
static String |
ITALICS |
static int |
ITALICS_ID |
static int |
NUM_ID |
static String |
SUB_HEADING |
static int |
SUB_HEADING_ID |
static String[] |
TOKEN_TYPES
String token types that correspond to token type int constants
|
static int |
TOKENS_ONLY
Only output tokens
|
static int |
UNTOKENIZED_ONLY
Only output untokenized tokens, which are tokens that would normally be split into several tokens
|
static int |
UNTOKENIZED_TOKEN_FLAG
This flag is used to indicate that the produced "Token" would, if
TOKENS_ONLY was used, produce multiple tokens. |
DEFAULT_TOKEN_ATTRIBUTE_FACTORY| Constructor and Description |
|---|
WikipediaTokenizer()
Creates a new instance of the
WikipediaTokenizer. |
WikipediaTokenizer(AttributeFactory factory,
int tokenOutput,
Set<String> untokenizedTypes)
Creates a new instance of the
WikipediaTokenizer. |
WikipediaTokenizer(int tokenOutput,
Set<String> untokenizedTypes)
Creates a new instance of the
WikipediaTokenizer. |
| Modifier and Type | Method and Description |
|---|---|
void |
close() |
void |
end() |
boolean |
incrementToken() |
void |
reset() |
correctOffset, setReaderaddAttribute, addAttributeImpl, captureState, clearAttributes, cloneAttributes, copyTo, endAttributes, equals, getAttribute, getAttributeClassesIterator, getAttributeFactory, getAttributeImplsIterator, hasAttribute, hasAttributes, hashCode, reflectAsString, reflectWith, removeAllAttributes, restoreState, toStringpublic static final String INTERNAL_LINK
public static final String EXTERNAL_LINK
public static final String EXTERNAL_LINK_URL
public static final String CITATION
public static final String CATEGORY
public static final String BOLD
public static final String ITALICS
public static final String BOLD_ITALICS
public static final String HEADING
public static final String SUB_HEADING
public static final int ALPHANUM_ID
public static final int APOSTROPHE_ID
public static final int ACRONYM_ID
public static final int COMPANY_ID
public static final int EMAIL_ID
public static final int HOST_ID
public static final int NUM_ID
public static final int CJ_ID
public static final int INTERNAL_LINK_ID
public static final int EXTERNAL_LINK_ID
public static final int CITATION_ID
public static final int CATEGORY_ID
public static final int BOLD_ID
public static final int ITALICS_ID
public static final int BOLD_ITALICS_ID
public static final int HEADING_ID
public static final int SUB_HEADING_ID
public static final int EXTERNAL_LINK_URL_ID
public static final String[] TOKEN_TYPES
public static final int TOKENS_ONLY
public static final int UNTOKENIZED_ONLY
public static final int BOTH
public static final int UNTOKENIZED_TOKEN_FLAG
TOKENS_ONLY was used, produce multiple tokens.public WikipediaTokenizer()
WikipediaTokenizer. Attaches the
input to a newly created JFlex scanner.public WikipediaTokenizer(int tokenOutput,
Set<String> untokenizedTypes)
WikipediaTokenizer. Attaches the
input to the newly created JFlex scanner.tokenOutput - One of TOKENS_ONLY, UNTOKENIZED_ONLY, BOTHpublic WikipediaTokenizer(AttributeFactory factory, int tokenOutput, Set<String> untokenizedTypes)
WikipediaTokenizer. Attaches the
input to the newly created JFlex scanner. Uses the given AttributeFactory.tokenOutput - One of TOKENS_ONLY, UNTOKENIZED_ONLY, BOTHpublic final boolean incrementToken()
throws IOException
incrementToken in class TokenStreamIOExceptionpublic void close()
throws IOException
close in interface Closeableclose in interface AutoCloseableclose in class TokenizerIOExceptionpublic void reset()
throws IOException
reset in class TokenizerIOExceptionpublic void end()
throws IOException
end in class TokenStreamIOExceptionCopyright © 2000-2021 Apache Software Foundation. All Rights Reserved.