GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism | allainews.com

June 26, 2024, 4:45 a.m. | Byungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim, Sunghyun Park, Neeraj Aggarwal, Colin Unger, Daiyaan Arfeen, Peiyuan Liao, Xupeng Miao, Mohammad Al

cs.LG updates on arXiv.org arxiv.org

arXiv:2406.17145v1 Announce Type: cross
Abstract: Deep neural networks (DNNs) continue to grow rapidly in size, making them infeasible to train on a single device. Pipeline parallelism is commonly used in existing DNN systems to support large-scale DNN training by partitioning a DNN into multiple stages, which concurrently perform DNN training for different micro-batches in a pipeline fashion. However, existing pipeline-parallel approaches only consider sequential pipeline stages and thus ignore the topology of a DNN, resulting in missed model-parallel opportunities. This …

abstract arxiv cs.ai cs.dc cs.lg device dnn graph grow improving making multiple networks neural networks partitioning performance pipeline scalability scale stages support systems them train training type

More from arxiv.org / cs.LG updates on arXiv.org

Bayesian identification of nonseparable Hamiltonians with multiplicative noise using deep learning and reduced-order modeling 2 days, 16 hours ago | arxiv.org

abstract arxiv bayesian cs.lg +17

MMGPL: Multimodal Medical Data Analysis with Graph Prompt Learning 2 days, 16 hours ago | arxiv.org

abstract analysis arxiv cs.cv +16

Self-Supervised Detection of Perfect and Partial Input-Dependent Symmetries 2 days, 16 hours ago | arxiv.org

arxiv cs.cv cs.lg detection +3

MixerFlow: MLP-Mixer meets Normalising Flows 2 days, 16 hours ago | arxiv.org

abstract architectures arxiv context +15

Machine Learning-Enabled Software and System Architecture Frameworks 2 days, 16 hours ago | arxiv.org

abstract architecture arxiv concerns +22

Efficient Interaction-Aware Interval Analysis of Neural Network Feedback Loops 2 days, 16 hours ago | arxiv.org

abstract analysis arxiv cs.lg +19

Kernelised Normalising Flows 2 days, 16 hours ago | arxiv.org

abstract architecture arxiv capabilities +14

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism 2 days, 16 hours ago | arxiv.org

abstract arxiv class cs.dc +25

Reinforcement Learning in Credit Scoring and Underwriting 2 days, 16 hours ago | arxiv.org

abstract action adapt arxiv +17

Junior Senior Reliability Engineer

@ NielsenIQ | Bogotá, Colombia

View on ai-jobs.net

[Job - 15712] Vaga Afirmativa para Mulheres - QA (Automation), SR

@ CI&T | Brazil

View on ai-jobs.net

Production Reliability Engineer, Trade Desk

@ Jump Trading | Sydney, Australia

View on ai-jobs.net

Senior Process Engineer, Prenatal

@ BillionToOne | Union City and Menlo Park, CA

View on ai-jobs.net

Senior Scientist, Sustainability Science and Innovation

@ Microsoft | Redmond, Washington, United States

View on ai-jobs.net

Data Scientist

@ Ford Motor Company | Chennai, Tamil Nadu, India

View on ai-jobs.net