Package org.apache.lucene.analysis.cjk
Analyzer for Chinese, Japanese, and Korean, which indexes bigrams (overlapping groups of two adjacent Han characters).
See:
Description
Class Summary |
CJKAnalyzer |
An Analyzer that tokenizes text with CJKTokenizer and
filters with StopFilter |
CJKTokenizer |
CJKTokenizer is designed for Chinese, Japanese, and Korean languages. |
Package org.apache.lucene.analysis.cjk Description
Analyzer for Chinese, Japanese, and Korean, which indexes bigrams (overlapping groups of two adjacent Han characters).
Three analyzers are provided for Chinese, each of which treats Chinese text in a different way.
- ChineseAnalyzer (in the analyzers/cn package): Index unigrams (individual Chinese characters) as a token.
- CJKAnalyzer (in this package): Index bigrams (overlapping groups of two adjacent Chinese characters) as tokens.
- SmartChineseAnalyzer (in the analyzers/smartcn package): Index words (attempt to segment Chinese text into words) as tokens.
Example phrase: "我是中国人"
- ChineseAnalyzer: 我-是-中-国-人
- CJKAnalyzer: 我是-是中-中国-国人
- SmartChineseAnalyzer: 我-是-中国-人
Copyright © 2000-2010 Apache Software Foundation. All Rights Reserved.