Simple Hack for Transformers against Heavy Long-Text Classification on a Time- and Memory-Limited GPU Service | allainews.com

March 20, 2024, 4:48 a.m. | Mirza Alim Mutasodirin, Radityo Eko Prasojo, Achmad F. Abka, Hanif Rasyidi

cs.CL updates on arXiv.org arxiv.org

arXiv:2403.12563v1 Announce Type: new
Abstract: Many NLP researchers rely on free computational services, such as Google Colab, to fine-tune their Transformer models, causing a limitation for hyperparameter optimization (HPO) in long-text classification due to the method having quadratic complexity and needing a bigger resource. In Indonesian, only a few works were found on long-text classification using Transformers. Most only use a small amount of data and do not report any HPO. In this study, using 18k news articles, we investigate …

abstract arxiv bigger classification colab complexity computational cs.ai cs.cl free google gpu hack hyperparameter memory nlp optimization researchers service services simple text text classification transformer transformer models transformers type

More from arxiv.org / cs.CL updates on arXiv.org

Hijacking Context in Large Multi-modal Models 6 hours ago | arxiv.org

abstract arxiv contents context +16

The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks 6 hours ago | arxiv.org

abstract arxiv concerns cs.cl +21

Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias 6 hours ago | arxiv.org

abstract accuracy agent arxiv +28

Small Language Model Can Self-correct 6 hours ago | arxiv.org

abstract arxiv capability chatgpt +18

Prompt-based mental health screening from social media text 6 hours ago | arxiv.org

abstract article arxiv bag +17

Scaling Political Texts with Large Language Models: Asking a Chatbot Might Be All You Need 6 hours ago | arxiv.org

abstract arxiv author chatbot +20

Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis 6 hours ago | arxiv.org

analysis arxiv attribution bias +10

Natural Language Interfaces for Tabular Data Querying and Visualization: A Survey 6 hours ago | arxiv.org

abstract arxiv chatgpt cs.ai +27

Hidden Citations Obscure True Impact in Science 6 hours ago | arxiv.org

abstract arxiv citations clear +19

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net

Research Engineer

@ Allora Labs | Remote

View on ai-jobs.net

Ecosystem Manager

@ Allora Labs | Remote

View on ai-jobs.net

Founding AI Engineer, Agents

@ Occam AI | New York

View on ai-jobs.net