我正在遵循谷歌云ml上重新训练《盗梦空间》的花儿教程。我可以运行教程,训练,预测,就可以了。在
然后我用flowers数据集替换了我自己的测试数据集。图像数字的光学字符识别。在
我的完整代码是here
labels的Dict文件
评估set
训练Set
运行于最近由谷歌提供的docker构建。在
`docker run -it -p "127.0.0.1:8080:8080" --entrypoint=/bin/bash gcr.io/cloud-datalab/datalab:local-20161227
我可以预处理文件,并提交培训作业使用
^{pr2}$但它从来没有通过全局步骤0。花教程大约1小时在免费层进行培训。我放弃了长达11小时的训练。不许动。在
看着stackdriver,什么进展都没有。在
我也尝试了一个由20个训练图像和10个评估图像组成的小玩具数据集。同样的问题。在
也许不足为奇的是,我无法在tensorboard上看到这个日志,没有什么可以显示的。在
完整的培训日志:
INFO 2017-01-10 17:22:00 +0000 unknown_task Validating job requirements...
INFO 2017-01-10 17:22:01 +0000 unknown_task Job creation request has been successfully validated.
INFO 2017-01-10 17:22:01 +0000 unknown_task Job MeerkatReader_MeerkatReader_20170110_170701 is queued.
INFO 2017-01-10 17:22:07 +0000 unknown_task Waiting for job to be provisioned.
INFO 2017-01-10 17:22:07 +0000 unknown_task Waiting for TensorFlow to start.
INFO 2017-01-10 17:22:10 +0000 master-replica-0 Running task with arguments: --cluster={"master": ["master-d4f6-0:2222"]} --task={"type": "master", "index": 0} --job={
INFO 2017-01-10 17:22:10 +0000 master-replica-0 "package_uris": ["gs://api-project-773889352370-ml/MeerkatReader_MeerkatReader_20170110_170701/f78d90a60f615a2d108d06557818eb4f82ffa94a/trainer-0.1.tar.gz"],
INFO 2017-01-10 17:22:10 +0000 master-replica-0 "python_module": "trainer.task",
INFO 2017-01-10 17:22:10 +0000 master-replica-0 "args": ["--output_path", "gs://api-project-773889352370-ml/MeerkatReader/MeerkatReader_MeerkatReader_20170110_170701/training", "--eval_data_paths", "gs://api-project-773889352370-ml/MeerkatReader/MeerkatReader_MeerkatReader_20170110_170701/preproc/eval*", "--train_data_paths", "gs://api-project-773889352370-ml/MeerkatReader/MeerkatReader_MeerkatReader_20170110_170701/preproc/train*"],
INFO 2017-01-10 17:22:10 +0000 master-replica-0 "region": "us-central1"
INFO 2017-01-10 17:22:10 +0000 master-replica-0 } --beta
INFO 2017-01-10 17:22:10 +0000 master-replica-0 Downloading the package: gs://api-project-773889352370-ml/MeerkatReader_MeerkatReader_20170110_170701/f78d90a60f615a2d108d06557818eb4f82ffa94a/trainer-0.1.tar.gz
INFO 2017-01-10 17:22:10 +0000 master-replica-0 Running command: gsutil -q cp gs://api-project-773889352370-ml/MeerkatReader_MeerkatReader_20170110_170701/f78d90a60f615a2d108d06557818eb4f82ffa94a/trainer-0.1.tar.gz trainer-0.1.tar.gz
INFO 2017-01-10 17:22:12 +0000 master-replica-0 Building wheels for collected packages: trainer
INFO 2017-01-10 17:22:12 +0000 master-replica-0 creating '/tmp/tmpSgdSzOpip-wheel-/trainer-0.1-cp27-none-any.whl' and adding '.' to it
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer/model.py'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer/util.py'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer/preprocess.py'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer/task.py'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer-0.1.dist-info/metadata.json'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer-0.1.dist-info/WHEEL'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 adding 'trainer-0.1.dist-info/METADATA'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 Running setup.py bdist_wheel for trainer: finished with status 'done'
INFO 2017-01-10 17:22:12 +0000 master-replica-0 Stored in directory: /root/.cache/pip/wheels/e8/0c/c7/b77d64796dbbac82503870c4881d606fa27e63942e07c75f0e
INFO 2017-01-10 17:22:12 +0000 master-replica-0 Successfully built trainer
INFO 2017-01-10 17:22:13 +0000 master-replica-0 Running command: python -m trainer.task --output_path gs://api-project-773889352370-ml/MeerkatReader/MeerkatReader_MeerkatReader_20170110_170701/training --eval_data_paths gs://api-project-773889352370-ml/MeerkatReader/MeerkatReader_MeerkatReader_20170110_170701/preproc/eval* --train_data_paths gs://api-project-773889352370-ml/MeerkatReader/MeerkatReader_MeerkatReader_20170110_170701/preproc/train*
INFO 2017-01-10 17:22:14 +0000 master-replica-0 Starting master/0
INFO 2017-01-10 17:22:14 +0000 master-replica-0 Initialize GrpcChannelCache for job master -> {0 -> localhost:2222}
INFO 2017-01-10 17:22:14 +0000 master-replica-0 Started server with target: grpc://localhost:2222
ERROR 2017-01-10 17:22:16 +0000 master-replica-0 device_filters: "/job:ps"
INFO 2017-01-10 17:22:19 +0000 master-replica-0 global_step/sec: 0
重复最后一行直到我把它弄死。在
我对这项服务的心理模型不正确吗?欢迎所有建议。在
一切看起来都很好。我怀疑你的数据有问题。我特别怀疑TF无法从你的GCS文件中读取任何数据(它们是空的吗?)?结果,当您调用train时,TF最终阻塞了它无法读取的一批数据。在
我建议在会话.运行在Trainer.run_training中。这会告诉你这条线是不是卡住了。在
我还建议检查一下你的GCS文件的大小。在
TensorFlow还有一个实验性的RunOptions,它允许您指定会话.运行. 一旦这个特性准备好了,这可能有助于确保代码不会永远阻塞。在
相关问题 更多 >
编程相关推荐