all AI news
Huggingface not saving model checkpoint
April 28, 2023, 1:37 p.m. | /u/Tiny-Entertainer-346
Natural Language Processing www.reddit.com
args = Seq2SeqTrainingArguments(
model_dir,
evaluation_strategy="steps",
eval_steps=100,
logging_strategy="steps",
logging_steps=100,
save_strategy="steps",
save_steps=200,
learning_rate=4e-5,
per_device_train_batch_size=batch_size,
per_device_eval_batch_size=batch_size,
weight_decay=0.01,
save_total_limit=3,
num_train_epochs=10,
predict_with_generate=True,
fp16=True,
load_best_model_at_end=True,
metric_for_best_model="rouge1",
report_to="tensorboard"
)
My model trained for 7600 steps. But the last model saved was for checkpoint 1800:
[trainer screenshot](https://i.stack.imgur.com/MBoFu.png)
Why is this so?
fp16 huggingface languagetechnology look saving tensorboard training true
More from www.reddit.com / Natural Language Processing
Which NLP-master programs in Europe are more cs-leaning?
2 days, 10 hours ago |
www.reddit.com
What do you think is the state of the art technique for matching a piece …
4 days, 8 hours ago |
www.reddit.com
Multilabel text classification on unlabled data
4 days, 21 hours ago |
www.reddit.com
Did we just receive an AI-generated meta-review?
1 week, 2 days ago |
www.reddit.com
Found a Way to Keep Transcripts Going 24/7
1 week, 3 days ago |
www.reddit.com
Jobs in AI, ML, Big Data
Founding AI Engineer, Agents
@ Occam AI | New York
AI Engineer Intern, Agents
@ Occam AI | US
AI Research Scientist
@ Vara | Berlin, Germany and Remote
Data Architect
@ University of Texas at Austin | Austin, TX
Data ETL Engineer
@ University of Texas at Austin | Austin, TX
Machine Learning Engineer
@ Apple | Sunnyvale, California, United States