【Bug已解决】PyTorch: RuntimeError: Input, output and indices must be on the current device 解决方案 【Bug已解决】PyTorch: RuntimeError: Input, output and indices must be on the current device 解决方案问题描述在 PyTorch 中进行深度学习模型训练或推理时当你尝试将模型、输入数据、损失函数或优化器分布在不同设备CPU 与 GPU或多个 GPU上时经常会遇到如下错误RuntimeError: Input, output and indices must be on the current device这个错误通常出现在使用torch.nn.NLLLoss、torch.nn.CrossEntropyLoss、torch.nn.Embedding或其他需要索引操作的模块时。其核心含义是参与计算的张量tensor不在同一个设备上导致 PyTorch 无法跨设备执行操作。该问题在以下场景中尤为常见初次将模型迁移到 GPU 时忘记同步迁移标签label张量使用多 GPU 训练DataParallel或DistributedDataParallel时设备分配不一致在混合精度训练AMP中部分张量被自动转移到 GPU 而部分仍留在 CPU使用预训练模型时模型权重在 GPU 上但输入数据未转移错误复现以下代码可以稳定复现该错误import torch import torch.nn as nn # 定义一个简单的文本分类模型 class TextClassifier(nn.Module): def __init__(self, vocab_size, embed_dim, num_classes): super(TextClassifier, self).__init__() self.embedding nn.Embedding(vocab_size, embed_dim) self.fc nn.Linear(embed_dim, num_classes) def forward(self, x): # x 是整数索引张量用于 Embedding 查表 embedded self.embedding(x) # Embedding 需要 indices 在同一设备 pooled embedded.mean(dim1) return self.fc(pooled) # 模型参数 vocab_size 10000 embed_dim 128 num_classes 5 batch_size 32 seq_len 20 # 创建模型并移到 GPU device torch.device(cuda:0) model TextClassifier(vocab_size, embed_dim, num_classes).to(device) # 创建模拟输入数据整数索引和标签 # 故意将输入数据留在 CPU 上不转移到 GPU input_data torch.randint(0, vocab_size, (batch_size, seq_len)) # 在 CPU 上 labels torch.randint(0, num_classes, (batch_size,)) # 在 CPU 上 # 损失函数 criterion nn.CrossEntropyLoss() # 前向传播 —— 这里会报错 # 因为 model 在 GPU 上embedding 权重在 GPU 上 # 但 input_dataindices在 CPU 上 try: outputs model(input_data) loss criterion(outputs, labels) except RuntimeError as e: print(f错误信息: {e}) # 输出: RuntimeError: Input, output and indices must be on the current device运行上述代码后你会在self.embedding(x)这一行看到错误。原因是nn.Embedding的权重矩阵在 GPU 上但输入索引x还在 CPU 上PyTorch 不允许跨设备索引操作。根因分析1. PyTorch 的设备一致性要求PyTorch 中的每个张量都有一个device属性标明它存储在 CPU 还是某个 GPU 上。当执行运算时PyTorch 要求所有参与运算的张量必须在同一设备上。对于普通算术运算加、减、乘等如果张量在不同设备上PyTorch 通常会给出明确的设备不匹配错误。但对于索引类操作如Embedding查表、NLLLoss计算等错误信息就是 Input, output and indices must be on the current device。2..to(device)只转移了模型权重调用model.to(device)只会将模型内部的参数weights/biases和缓冲区buffers转移到目标设备但不会自动转移你传入的输入数据。这是最常见的误解来源model.to(device) # 只转移模型参数到 GPU input_data.to(device) # 需要单独转移输入数据 labels.to(device) # 需要单独转移标签3. Embedding 层的特殊性nn.Embedding本质上是一个查表操作给定一组整数索引从权重矩阵中取出对应的行。这个操作要求索引张量和权重矩阵在同一设备上。如果权重在 GPU 而索引在 CPUGPU 无法直接访问 CPU 内存中的索引数据反之亦然。4. 损失函数的设备要求CrossEntropyLoss内部包含LogSoftmax和NLLLoss其中NLLLoss需要用标签索引模型输出。如果模型输出在 GPU 上但标签在 CPU 上同样会触发此错误。解决方案方案一统一转移所有张量到同一设备推荐最直接、最可靠的方案是确保模型、输入数据、标签都在同一设备上import torch import torch.nn as nn class TextClassifier(nn.Module): def __init__(self, vocab_size, embed_dim, num_classes): super(TextClassifier, self).__init__() self.embedding nn.Embedding(vocab_size, embed_dim) self.fc nn.Linear(embed_dim, num_classes) def forward(self, x): embedded self.embedding(x) pooled embedded.mean(dim1) return self.fc(pooled) # 修复后的完整训练代码 vocab_size 10000 embed_dim 128 num_classes 5 batch_size 32 seq_len 20 num_epochs 5 learning_rate 0.001 # 确定设备 device torch.device(cuda:0 if torch.cuda.is_available() else cpu) print(f使用设备: {device}) # 创建模型并转移到设备 model TextClassifier(vocab_size, embed_dim, num_classes).to(device) # 创建模拟数据 input_data torch.randint(0, vocab_size, (batch_size, seq_len)) labels torch.randint(0, num_classes, (batch_size,)) # 关键修复将输入数据和标签都转移到同一设备 input_data input_data.to(device) labels labels.to(device) # 损失函数和优化器 criterion nn.CrossEntropyLoss() optimizer torch.optim.Adam(model.parameters(), lrlearning_rate) # 训练循环 for epoch in range(num_epochs): model.train() optimizer.zero_grad() # 前向传播 outputs model(input_data) loss criterion(outputs, labels) # 反向传播 loss.backward() optimizer.step() print(fEpoch [{epoch1}/{num_epochs}], Loss: {loss.item():.4f}) print(训练完成)方案二封装设备转移逻辑在实际项目中建议封装一个通用的设备转移函数避免遗漏import torch import torch.nn as nn from torch.utils.data import Dataset, DataLoader class TextDataset(Dataset): 自定义数据集 def __init__(self, num_samples, vocab_size, seq_len, num_classes): self.data torch.randint(0, vocab_size, (num_samples, seq_len)) self.labels torch.randint(0, num_classes, (num_samples,)) def __len__(self): return len(self.labels) def __getitem__(self, idx): return self.data[idx], self.labels[idx] def get_default_device(): 获取默认设备 return torch.device(cuda:0 if torch.cuda.is_available() else cpu) def to_device(data, device): 递归地将数据转移到指定设备 if isinstance(data, (list, tuple)): return [to_device(x, device) for x in data] elif isinstance(data, dict): return {k: to_device(v, device) for k, v in data.items()} else: return data.to(device, non_blockingTrue) class DeviceDataLoader: 包装 DataLoader自动将每个 batch 转移到 GPU def __init__(self, dataloader, device): self.dataloader dataloader self.device device def __iter__(self): for batch in self.dataloader: yield to_device(batch, self.device) def __len__(self): return len(self.dataloader) # 使用示例 vocab_size 10000 embed_dim 128 num_classes 5 seq_len 20 batch_size 32 num_epochs 5 device get_default_device() # 创建数据集和 DataLoader dataset TextDataset(1000, vocab_size, seq_len, num_classes) train_loader DataLoader(dataset, batch_sizebatch_size, shuffleTrue) # 用 DeviceDataLoader 包装自动转移设备 train_loader DeviceDataLoader(train_loader, device) # 创建模型 model TextClassifier(vocab_size, embed_dim, num_classes).to(device) # 训练 criterion nn.CrossEntropyLoss() optimizer torch.optim.Adam(model.parameters(), lr0.001) for epoch in range(num_epochs): model.train() total_loss 0 for batch_inputs, batch_labels in train_loader: # batch_inputs 和 batch_labels 已经在 GPU 上了 optimizer.zero_grad() outputs model(batch_inputs) loss criterion(outputs, batch_labels) loss.backward() optimizer.step() total_loss loss.item() avg_loss total_loss / len(train_loader) print(fEpoch [{epoch1}/{num_epochs}], Avg Loss: {avg_loss:.4f}) print(训练完成)方案三在 forward 方法中自动转移有时你希望模型能够自动处理设备不一致的情况可以在forward方法中添加设备同步逻辑class RobustTextClassifier(nn.Module): def __init__(self, vocab_size, embed_dim, num_classes): super().__init__() self.embedding nn.Embedding(vocab_size, embed_dim) self.fc nn.Linear(embed_dim, num_classes) def forward(self, x): # 自动将输入转移到与模型参数相同的设备 device next(self.parameters()).device if x.device ! device: x x.to(device) embedded self.embedding(x) pooled embedded.mean(dim1) return self.fc(pooled) class RobustTrainer: 封装训练逻辑自动处理设备转移 def __init__(self, model, lr0.001): self.device torch.device(cuda:0 if torch.cuda.is_available() else cpu) self.model model.to(self.device) self.criterion nn.CrossEntropyLoss() self.optimizer torch.optim.Adam(model.parameters(), lrlr) def train_step(self, inputs, labels): self.model.train() self.optimizer.zero_grad() # 模型内部会自动转移 inputs outputs self.model(inputs) # 标签需要手动转移因为 criterion 在外部 labels labels.to(self.device) loss self.criterion(outputs, labels) loss.backward() self.optimizer.step() return loss.item() def evaluate(self, inputs, labels): self.model.eval() with torch.no_grad(): outputs self.model(inputs) labels labels.to(self.device) loss self.criterion(outputs, labels) _, predicted torch.max(outputs, 1) accuracy (predicted labels).float().mean() return loss.item(), accuracy.item() # 使用 model RobustTextClassifier(10000, 128, 5) trainer RobustTrainer(model) for epoch in range(5): inputs torch.randint(0, 10000, (32, 20)) labels torch.randint(0, 5, (32,)) loss trainer.train_step(inputs, labels) print(fEpoch {epoch1}, Loss: {loss:.4f})完整修复代码以下是一个完整的、可直接运行的端到端修复示例包含数据加载、模型定义、训练和评估import torch import torch.nn as nn from torch.utils.data import Dataset, DataLoader import numpy as np # 数据集定义 class TextClassificationDataset(Dataset): 模拟文本分类数据集 def __init__(self, num_samples1000, vocab_size10000, seq_len20, num_classes5): self.data torch.randint(0, vocab_size, (num_samples, seq_len), dtypetorch.long) self.labels torch.randint(0, num_classes, (num_samples,), dtypetorch.long) def __len__(self): return len(self.labels) def __getitem__(self, idx): return self.data[idx], self.labels[idx] # 模型定义 class TextClassifier(nn.Module): def __init__(self, vocab_size, embed_dim, hidden_dim, num_classes, dropout0.3): super().__init__() self.embedding nn.Embedding(vocab_size, embed_dim, padding_idx0) self.lstm nn.LSTM(embed_dim, hidden_dim, batch_firstTrue, num_layers2, dropoutdropout) self.dropout nn.Dropout(dropout) self.fc nn.Linear(hidden_dim, num_classes) def forward(self, x): embedded self.embedding(x) lstm_out, (hidden, _) self.lstm(embedded) # 使用最后一层的隐藏状态 hidden hidden[-1] output self.dropout(hidden) return self.fc(output) # 设备管理工具 class DeviceManager: 集中管理设备转移避免设备不一致 def __init__(self): self.device torch.device(cuda:0 if torch.cuda.is_available() else cpu) def to_device(self, *tensors): 将多个张量转移到当前设备 return [t.to(self.device) for t in tensors] def move_model(self, model): 将模型转移到当前设备 return model.to(self.device) def __str__(self): return str(self.device) # 训练器 class Trainer: def __init__(self, model, train_loader, val_loader, lr0.001): self.dm DeviceManager() print(f训练设备: {self.dm}) self.model self.dm.move_model(model) self.train_loader train_loader self.val_loader val_loader self.criterion nn.CrossEntropyLoss() self.optimizer torch.optim.Adam(model.parameters(), lrlr) self.scheduler torch.optim.lr_scheduler.StepLR(self.optimizer, step_size3, gamma0.5) def train_epoch(self): self.model.train() total_loss 0 correct 0 total 0 for inputs, labels in self.train_loader: # 关键将输入和标签都转移到同一设备 inputs, labels self.dm.to_device(inputs, labels) self.optimizer.zero_grad() outputs self.model(inputs) loss self.criterion(outputs, labels) loss.backward() # 梯度裁剪 torch.nn.utils.clip_grad_norm_(self.model.parameters(), max_norm1.0) self.optimizer.step() total_loss loss.item() _, predicted outputs.max(1) correct (predicted labels).sum().item() total labels.size(0) return total_loss / len(self.train_loader), correct / total def validate(self): self.model.eval() total_loss 0 correct 0 total 0 with torch.no_grad(): for inputs, labels in self.val_loader: inputs, labels self.dm.to_device(inputs, labels) outputs self.model(inputs) loss self.criterion(outputs, labels) total_loss loss.item() _, predicted outputs.max(1) correct (predicted labels).sum().item() total labels.size(0) return total_loss / len(self.val_loader), correct / total def train(self, num_epochs10): best_val_acc 0 for epoch in range(num_epochs): train_loss, train_acc self.train_epoch() val_loss, val_acc self.validate() self.scheduler.step() print(fEpoch [{epoch1}/{num_epochs}]) print(f Train Loss: {train_loss:.4f}, Train Acc: {train_acc:.4f}) print(f Val Loss: {val_loss:.4f}, Val Acc: {val_acc:.4f}) if val_acc best_val_acc: best_val_acc val_acc torch.save(self.model.state_dict(), best_model.pth) print(f - 保存最佳模型 (val_acc{val_acc:.4f})) print(f\n训练完成最佳验证准确率: {best_val_acc:.4f}) # 主程序 if __name__ __main__: # 超参数 VOCAB_SIZE 10000 EMBED_DIM 128 HIDDEN_DIM 256 NUM_CLASSES 5 BATCH_SIZE 64 NUM_EPOCHS 10 LR 0.001 # 创建数据集 train_dataset TextClassificationDataset(800, VOCAB_SIZE, 20, NUM_CLASSES) val_dataset TextClassificationDataset(200, VOCAB_SIZE, 20, NUM_CLASSES) train_loader DataLoader(train_dataset, batch_sizeBATCH_SIZE, shuffleTrue) val_loader DataLoader(val_dataset, batch_sizeBATCH_SIZE, shuffleFalse) # 创建模型 model TextClassifier(VOCAB_SIZE, EMBED_DIM, HIDDEN_DIM, NUM_CLASSES) # 训练 trainer Trainer(model, train_loader, val_loader, lrLR) trainer.train(num_epochsNUM_EPOCHS) # 推理示例 dm DeviceManager() model model.to(dm.device) model.eval() test_input torch.randint(0, VOCAB_SIZE, (1, 20)) test_input test_input.to(dm.device) # 确保输入在正确设备上 with torch.no_grad(): output model(test_input) prediction output.argmax(dim1) print(f\n推理结果: 预测类别 {prediction.item()})常见陷阱与注意事项1. 忘记转移标签张量最常见的错误是只转移了输入数据但忘了转移标签# 错误写法 inputs inputs.to(device) outputs model(inputs) loss criterion(outputs, labels) # labels 还在 CPU 上 # 正确写法 inputs inputs.to(device) labels labels.to(device) outputs model(inputs) loss criterion(outputs, labels)2.DataParallel下的设备问题使用DataParallel时模型会在多个 GPU 上复制但输入数据会被自动分割。需要注意的是DataParallel要求输入在cuda:0上它会自动分发到其他 GPUmodel nn.DataParallel(model, device_ids[0, 1, 2]) model model.to(cuda:0) # DataParallel 要求主设备是 cuda:0 # 输入也必须在 cuda:0 上 inputs inputs.to(cuda:0) labels labels.to(cuda:0)3. 混合精度训练AMP中的设备问题使用torch.cuda.amp时autocast不会自动转移设备仍需手动确保一致性scaler torch.cuda.amp.GradScaler() for inputs, labels in train_loader: inputs inputs.to(device) labels labels.to(device) optimizer.zero_grad() with torch.cuda.amp.autocast(): outputs model(inputs) loss criterion(outputs, labels) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update()4. 模型保存和加载时的设备问题加载模型时如果不指定设备可能会将 GPU 模型加载到 CPU 上导致后续推理时设备不一致# 保存时 torch.save(model.state_dict(), model.pth) # 加载时指定设备 device torch.device(cuda:0 if torch.cuda.is_available() else cpu) model.load_state_dict(torch.load(model.pth, map_locationdevice)) model model.to(device)5. 自定义层中的设备不一致如果你在自定义层中创建了张量如常量矩阵这些张量默认在 CPU 上不会随model.to(device)转移。必须使用register_buffer注册class MyLayer(nn.Module): def __init__(self, dim): super().__init__() # 错误这个张量不会随 model.to(device) 转移 # self.mask torch.ones(dim) # 正确使用 register_buffer self.register_buffer(mask, torch.ones(dim)) def forward(self, x): return x * self.mask6. 使用non_blockingTrue的注意事项当使用pin_memoryTrue的 DataLoader 时可以使用non_blockingTrue来实现异步数据传输提高效率# DataLoader 使用 pin_memory train_loader DataLoader(dataset, batch_size64, shuffleTrue, pin_memoryTrue) # 转移时使用 non_blocking for inputs, labels in train_loader: inputs inputs.to(device, non_blockingTrue) labels labels.to(device, non_blockingTrue) # ...但注意non_blockingTrue只在源张量是 pinned memory 时才真正异步否则等同于同步操作。总结RuntimeError: Input, output and indices must be on the current device是 PyTorch 中最常见的设备不一致错误之一。其根本原因是参与计算的张量分布在不同的设备CPU/GPU上而 PyTorch 不支持跨设备的索引操作。解决该问题的核心原则是确保模型参数、输入数据、标签张量以及损失函数中涉及的所有张量都在同一设备上。具体做法包括手动转移在每次前向传播前使用.to(device)将输入和标签转移到模型所在设备封装转移逻辑使用DeviceDataLoader或DeviceManager等工具类自动处理设备转移在 forward 中自动同步通过next(self.parameters()).device获取模型设备并自动转移输入使用register_buffer确保自定义层中的常量张量能随模型一起转移设备通过建立统一的设备管理策略可以从根本上避免此类错误使训练代码更加健壮和可维护。在实际项目中推荐使用方案二的封装方式将设备转移逻辑集中管理既减少了出错的可能性也提高了代码的可读性。