kuromoji_part_of_speech token filter
The kuromoji_part_of_speech
token filter removes tokens that match a set of part-of-speech tags. It accepts the following setting:
stoptags
- An array of part-of-speech tags that should be removed. It defaults to the
stoptags.txt
file embedded in thelucene-analyzer-kuromoji.jar
.
For example:
PUT kuromoji_sample
{
"settings": {
"index": {
"analysis": {
"analyzer": {
"my_analyzer": {
"tokenizer": "kuromoji_tokenizer",
"filter": [
"my_posfilter"
]
}
},
"filter": {
"my_posfilter": {
"type": "kuromoji_part_of_speech",
"stoptags": [
"ε©θ©-ζ Όε©θ©-δΈθ¬",
"ε©θ©-η΅ε©θ©"
]
}
}
}
}
}
}
GET kuromoji_sample/_analyze
{
"analyzer": "my_analyzer",
"text": "ε―ΏεΈγγγγγγ"
}
Which responds with:
{
"tokens" : [ {
"token" : "ε―ΏεΈ",
"start_offset" : 0,
"end_offset" : 2,
"type" : "word",
"position" : 0
}, {
"token" : "γγγγ",
"start_offset" : 3,
"end_offset" : 7,
"type" : "word",
"position" : 2
} ]
}