Active : 1 Machines / 4 Threads / 3.4 MNPS
Priority 0
RosentWinterquantize_net *diff40.0+0.40LLR: 1.31 (-2.94, 2.94) [0.00, 5.00]
Games: 716 W: 159 L: 105 D: 452
Ptnml(0-2): 3, 67, 171, 107, 10
Quantized network. Should be faster, at the cost of a little accuracy.
Priority -1
RosentWinternet382_d24f64b *diff8.0+0.08LLR: 0.48 (-2.94, 2.94) [0.00, 5.00]
Games: 8438 W: 2055 L: 1995 D: 4388
Ptnml(0-2): 181, 983, 1809, 1087, 159
Added net definition
Finished
RosentWinterquantize_net *diff8.0+0.08LLR: 2.97 (-2.94, 2.94) [0.00, 5.00]
Games: 1012 W: 301 L: 160 D: 551
Ptnml(0-2): 10, 80, 211, 169, 36
Quantized network. Should be faster, at the cost of a little accuracy.
RosentWinternet382_W50a *diff8.0+0.08LLR: -2.96 (-2.94, 2.94) [0.00, 5.00]
Games: 4510 W: 1066 L: 1178 D: 2266
Ptnml(0-2): 109, 580, 971, 504, 91
Net trained with more uniform data. If this is not negative we are happy because it simplifies our life. If this is negative it is fine as it means we were doing something reasonably fine.
RosentWinternet347_d24f64 *diff40.0+0.40LLR: -2.95 (-2.94, 2.94) [0.00, 5.00]
Games: 4604 W: 872 L: 969 D: 2763
Ptnml(0-2): 60, 555, 1154, 488, 45
Another day another net.
RosentWinternet347_d24f64 *diffN=40000Elo: 25.02 +- 7.56 (95%) [N=4000]
Games: 4020 W: 1277 L: 988 D: 1755
Ptnml(0-2): 95, 414, 767, 575, 159
Preliminary fixed node test with a potential larger net.
RosentWinternet347_W50 *diff40.0+0.40LLR: 2.96 (-2.94, 2.94) [0.00, 5.00]
Games: 4838 W: 986 L: 852 D: 3000
Ptnml(0-2): 44, 487, 1244, 579, 65
Trained with more data.
RosentWinternet347_W50 *diff8.0+0.08LLR: 3.03 (-2.94, 2.94) [0.00, 5.00]
Games: 2440 W: 644 L: 499 D: 1297
Ptnml(0-2): 38, 241, 541, 338, 62
Trained with more data.
RosentWintermaster *diff40.0+0.40Elo: 26.89 +- 6.25 (95%) [N=4000]
Games: 4000 W: 1002 L: 693 D: 2305
Ptnml(0-2): 33, 354, 963, 571, 79
Regression Test vs v5.0
RosentWinternet311_ace501v2 *diff40.0+0.40LLR: 2.95 (-2.94, 2.94) [0.00, 5.00]
Games: 7672 W: 1608 L: 1462 D: 4602
Ptnml(0-2): 83, 791, 1949, 923, 90
This is also trained with WSD but more data. I think I need to train with a longer cooldown phase and have a run with the old scheduler which will hopefully finish this afternoon.
RosentWinternet311_ace501v3 *diff8.0+0.08LLR: -3.05 (-2.94, 2.94) [0.00, 5.00]
Games: 29654 W: 7302 L: 7301 D: 15051
Ptnml(0-2): 656, 3723, 6123, 3614, 711
Comparison of old scheduler vs WSD. Based on my understanding the new scheduler is not properly tuned at the moment, so this should pass.
RosentWinternet311_ace501v2 *diff8.0+0.08LLR: 3.02 (-2.94, 2.94) [0.00, 5.00]
Games: 19066 W: 4895 L: 4667 D: 9504
Ptnml(0-2): 418, 2294, 3920, 2444, 457
This is also trained with WSD but more data. I think I need to train with a longer cooldown phase and have a run with the old scheduler which will hopefully finish this afternoon.
RosentWinternet311_ace501v2 *diff8.0+0.08LLR: -3.04 (-2.94, 2.94) [0.00, 5.00]
Games: 4132 W: 961 L: 1078 D: 2093
Ptnml(0-2): 111, 506, 918, 451, 80
Slight reduction in dataset size, but hopefully higher quality for training. Using WSD scheduler instead of linear. Hoping for small rating gain!
RosentWinternet311_ace501 *diff40.0+0.40LLR: 2.97 (-2.94, 2.94) [-5.00, 0.00]
Games: 15212 W: 2871 L: 2820 D: 9521
Ptnml(0-2): 109, 1621, 4104, 1654, 118
New approach to training net. Hopefully at least Elo neutral but less blunder prone.
RosentWinternet311_ace501 *diff8.0+0.08LLR: 2.95 (-2.94, 2.94) [-5.00, 0.00]
Games: 56412 W: 13466 L: 13594 D: 29352
Ptnml(0-2): 1179, 6842, 12275, 6748, 1162
New approach to training net. Hopefully at least Elo neutral but less blunder prone.
RosentWinternet_end_p *diff8.0+0.08LLR: -2.96 (-2.94, 2.94) [-5.00, 0.00]
Games: 20464 W: 5192 L: 5430 D: 9842
Ptnml(0-2): 599, 2484, 4207, 2440, 502
A much more reasonable middle ground.
RosentWinternet_end_p *diff8.0+0.08LLR: -2.98 (-2.94, 2.94) [-5.00, 0.00]
Games: 1216 W: 243 L: 393 D: 580
Ptnml(0-2): 54, 189, 236, 111, 18
Regression testing net training idea.
RosentWinternet_end_p *diff8.0+0.08LLR: 3.08 (-2.94, 2.94) [-5.00, 0.00]
Games: 13636 W: 3293 L: 3219 D: 7124
Ptnml(0-2): 272, 1627, 2966, 1661, 292
Non-regression sanity check.
RosentWinternet_end_pdiffN=64000Elo: 7.25 +- 13.90 (95%) [N=1000]
Games: 1006 W: 272 L: 251 D: 483
Ptnml(0-2): 19, 118, 213, 129, 24
Sanity test of net training
RosentWinternet_502_24x64 *diff8.0+0.08LLR: -2.96 (-2.94, 2.94) [0.00, 5.00]
Games: 6688 W: 1605 L: 1709 D: 3374
Ptnml(0-2): 168, 856, 1382, 788, 150
Warmup Stable Decay seems to fit the training pipeline better than the previous linear scheduler based on metrics. Does this hold in self play?
RosentWinternet_502_24x64 *diffN=40000Elo: 2.86 +- 7.28 (95%) [N=4000]
Games: 4006 W: 1111 L: 1078 D: 1817
Ptnml(0-2): 102, 478, 824, 483, 116
Get an understanding of how far apart the based increased size net is compared to the current version. Fixed nodes in this case as system load would otherwise distort results.
RosentWinternet_502_16x96 *diff10.0+0.10LLR: -2.95 (-2.94, 2.94) [-5.00, 0.00]
Games: 8948 W: 1999 L: 2170 D: 4779
Ptnml(0-2): 186, 1157, 1942, 1020, 169
The larger 24x64 net architecture is failing STC so we instead try 16x96. This has far fewer parameters but still has a 5% slowdown we need to overcome. Based on training metrics this lies squarely between the baseline and 24x64.
RosentWinternet_502_24x64 *diff10.0+0.10LLR: -2.95 (-2.94, 2.94) [-5.00, 0.00]
Games: 1566 W: 309 L: 453 D: 804
Ptnml(0-2): 65, 210, 333, 154, 21
Having passed the fixed node test as expected, we try an STC test with [-5, 0] bounds, as the slowdown is expected to hurt more than in actual play.
RosentWinternet_502_24x64 *diffN=40000LLR: 2.99 (-2.94, 2.94) [0.00, 5.00]
Games: 2282 W: 745 L: 582 D: 955
Ptnml(0-2): 69, 209, 455, 306, 102
Larger net, hopefully a clear improvement at fixed nodes. Around 16% less Nps.
RosentWinternet_311m *diff8.0+0.08LLR: -3.02 (-2.94, 2.94) [-2.50, 2.50]
Games: 17246 W: 4156 L: 4292 D: 8798
Ptnml(0-2): 402, 2127, 3657, 2079, 358
Trying to better understand the hybrid loss. Both nets are not fully trained, but are at a similar point in training.
RosentWinternet_311l *diff8.0+0.08LLR: -3.10 (-2.94, 2.94) [-3.00, 2.00]
Games: 17742 W: 4218 L: 4374 D: 9150
Ptnml(0-2): 424, 2152, 3826, 2094, 375
Trying to better understand the hybrid loss. Both nets are not fully trained, but are at a similar point in training. 311n removes the CE loss component completely. Even if worth Elo, we may not want this.
RosentWinternet_311l *diff8.0+0.08LLR: -2.97 (-2.94, 2.94) [0.00, 5.00]
Games: 23074 W: 5479 L: 5505 D: 12090
Ptnml(0-2): 487, 2821, 4936, 2817, 476
Final net with less CE loss component.