Tokenization Preference for Human and ML Model: An Annotation Study | allainews.com

Feb. 16, 2024, 5:48 a.m. | Tatsuya Hiraoka, Tomoya Iwakura

cs.CL updates on arXiv.org arxiv.org

arXiv:2304.10813v2 Announce Type: replace
Abstract: Is preferred tokenization for humans also preferred for machine-learning (ML) models? This study examines the relations between preferred tokenization for humans (appropriateness and readability) and one for ML models (performance on an NLP task). The question texts of the Japanese commonsense question-answering dataset are tokenized with six different tokenizers, and the performances of human annotators and ML models were compared. Furthermore, we analyze relations among performance of answers by human and ML model, the appropriateness …

abstract annotation arxiv cs.cl dataset human humans japanese machine ml models nlp performance question readability relations study tokenization type

More from arxiv.org / cs.CL updates on arXiv.org

STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models 1 day, 18 hours ago | arxiv.org

abstract arxiv computational cost +14

Large Language Models can Learn Rules 1 day, 18 hours ago | arxiv.org

abstract arxiv cs.ai cs.cl +18

Benchmarking LLMs via Uncertainty Quantification 1 day, 18 hours ago | arxiv.org

abstract arxiv benchmarking bridge +21

CARE: Extracting Experimental Findings From Clinical Literature 1 day, 18 hours ago | arxiv.org

abstract annotation applications arxiv +16

Prompt Cache: Modular Attention Reuse for Low-Latency Inference 1 day, 18 hours ago | arxiv.org

abstract arxiv attention cache +20

SpeechAlign: a Framework for Speech Translation Alignment Evaluation 1 day, 18 hours ago | arxiv.org

abstract advance alignment arxiv +14

I3: Intent-Introspective Retrieval Conditioned on Instructions 1 day, 18 hours ago | arxiv.org

abstract arxiv challenge cs.cl +10

DoDo Learning: DOmain-DemOgraphic Transfer in Language Models for Detecting Abuse Targeted at Public Figures 1 day, 18 hours ago | arxiv.org

abstract abuse arxiv automated +19

Investigating the prompt leakage effect and black-box defenses for multi-turn LLM interactions 1 day, 18 hours ago | arxiv.org

abstract arxiv box cs.ai +24

Data Architect

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Data ETL Engineer

@ University of Texas at Austin | Austin, TX

View on ai-jobs.net

Lead GNSS Data Scientist

@ Lurra Systems | Melbourne

View on ai-jobs.net

Senior Machine Learning Engineer (MLOps)

@ Promaton | Remote, Europe

View on ai-jobs.net

Director, Clinical Data Science

@ Aura | Remote USA

View on ai-jobs.net

Research Scientist, AI (PhD)

@ Meta | Menlo Park, CA | New York City

View on ai-jobs.net