Split Khmer sentences into word arrays using Seanghay Yath's split-khmer package, then compare the result against expected word boundaries. / បំបែកប្រយោគខ្មែរទៅជាបញ្ជីពាក្យដោយប្រើកញ្ចប់ split-khmer របស់ Seanghay Yath ហើយប្រៀបធៀបលទ្ធផលជាមួយព្រំដែនពាក្យដែលរំពឹងទុក។
កម្ពុជា មាន ភាសាខ្មែរ
[ "កម្ពុជា", "មាន", "ភាសាខ្មែរ" ]
MIT · original JavaScript splitter
កម្ពុជា មាន ភាសាខ្មែរ
F1: 80% · 2 matching boundaries / ព្រំដែនត្រូវគ្នា
Source · split-khmerMIT · deterministic maximum-matching toolkit
ក | ម្ពុ | ជា | មា | ន | ភាសា | ខ្មែរ
F1: 67% · 3 matching boundaries / ព្រំដែនត្រូវគ្នា
Source · khmer-nlp-toolkitThis is a Khmer Character Cluster view from the project's annotation/tooling approach, not a third word-segmentation model. KCC boundaries must not be interpreted as word boundaries. / នេះជាទិដ្ឋភាព Khmer Character Cluster តាមវិធីសាស្ត្រឧបករណ៍របស់គម្រោង មិនមែនជាម៉ូដែលបំបែកពាក្យទីបីទេ។ ព្រំដែន KCC មិនគួរបកស្រាយថាជាព្រំដែនពាក្យឡើយ។
ក | ម្ពុ | ជា | មា | ន | ភា | សា | ខ្មែ | រ
9 KCC clusters · Apache-2.0 project
Source · khmer-word-segmentationMatching boundaries / ព្រំដែនត្រូវគ្នា
2
Precision / ភាពត្រឹមត្រូវ
100%
Recall / អត្រារកឃើញ
67%
F1 score / ពិន្ទុ F1
80%