| Finished | ||||||
|---|---|---|---|---|---|---|
| Rosent | Winter | master * | diff | 40.0+0.40 | Elo: 26.89 +- 6.25 (95%) [N=4000] Games: 4000 W: 1002 L: 693 D: 2305 Ptnml(0-2): 33, 354, 963, 571, 79 | Regression Test vs v5.0 |
| Rosent | Winter | net311_ace501v2 * | diff | 40.0+0.40 | LLR: 2.95 (-2.94, 2.94) [0.00, 5.00] Games: 7672 W: 1608 L: 1462 D: 4602 Ptnml(0-2): 83, 791, 1949, 923, 90 | This is also trained with WSD but more data. I think I need to train with a longer cooldown phase and have a run with the old scheduler which will hopefully finish this afternoon. |
| Rosent | Winter | net311_ace501v3 * | diff | 8.0+0.08 | LLR: -3.05 (-2.94, 2.94) [0.00, 5.00] Games: 29654 W: 7302 L: 7301 D: 15051 Ptnml(0-2): 656, 3723, 6123, 3614, 711 | Comparison of old scheduler vs WSD. Based on my understanding the new scheduler is not properly tuned at the moment, so this should pass. |
| Rosent | Winter | net311_ace501v2 * | diff | 8.0+0.08 | LLR: 3.02 (-2.94, 2.94) [0.00, 5.00] Games: 19066 W: 4895 L: 4667 D: 9504 Ptnml(0-2): 418, 2294, 3920, 2444, 457 | This is also trained with WSD but more data. I think I need to train with a longer cooldown phase and have a run with the old scheduler which will hopefully finish this afternoon. |
| Rosent | Winter | net311_ace501v2 * | diff | 8.0+0.08 | LLR: -3.04 (-2.94, 2.94) [0.00, 5.00] Games: 4132 W: 961 L: 1078 D: 2093 Ptnml(0-2): 111, 506, 918, 451, 80 | Slight reduction in dataset size, but hopefully higher quality for training. Using WSD scheduler instead of linear. Hoping for small rating gain! |
| Rosent | Winter | net311_ace501 * | diff | 40.0+0.40 | LLR: 2.97 (-2.94, 2.94) [-5.00, 0.00] Games: 15212 W: 2871 L: 2820 D: 9521 Ptnml(0-2): 109, 1621, 4104, 1654, 118 | New approach to training net. Hopefully at least Elo neutral but less blunder prone. |
| Rosent | Winter | net311_ace501 * | diff | 8.0+0.08 | LLR: 2.95 (-2.94, 2.94) [-5.00, 0.00] Games: 56412 W: 13466 L: 13594 D: 29352 Ptnml(0-2): 1179, 6842, 12275, 6748, 1162 | New approach to training net. Hopefully at least Elo neutral but less blunder prone. |
| Rosent | Winter | net_end_p * | diff | 8.0+0.08 | LLR: -2.96 (-2.94, 2.94) [-5.00, 0.00] Games: 20464 W: 5192 L: 5430 D: 9842 Ptnml(0-2): 599, 2484, 4207, 2440, 502 | A much more reasonable middle ground. |
| Rosent | Winter | net_end_p * | diff | 8.0+0.08 | LLR: -2.98 (-2.94, 2.94) [-5.00, 0.00] Games: 1216 W: 243 L: 393 D: 580 Ptnml(0-2): 54, 189, 236, 111, 18 | Regression testing net training idea. |
| Rosent | Winter | net_end_p * | diff | 8.0+0.08 | LLR: 3.08 (-2.94, 2.94) [-5.00, 0.00] Games: 13636 W: 3293 L: 3219 D: 7124 Ptnml(0-2): 272, 1627, 2966, 1661, 292 | Non-regression sanity check. |
| Rosent | Winter | net_end_p | diff | N=64000 | Elo: 7.25 +- 13.90 (95%) [N=1000] Games: 1006 W: 272 L: 251 D: 483 Ptnml(0-2): 19, 118, 213, 129, 24 | Sanity test of net training |
| Rosent | Winter | net_502_24x64 * | diff | 8.0+0.08 | LLR: -2.96 (-2.94, 2.94) [0.00, 5.00] Games: 6688 W: 1605 L: 1709 D: 3374 Ptnml(0-2): 168, 856, 1382, 788, 150 | Warmup Stable Decay seems to fit the training pipeline better than the previous linear scheduler based on metrics. Does this hold in self play? |
| Rosent | Winter | net_502_24x64 * | diff | N=40000 | Elo: 2.86 +- 7.28 (95%) [N=4000] Games: 4006 W: 1111 L: 1078 D: 1817 Ptnml(0-2): 102, 478, 824, 483, 116 | Get an understanding of how far apart the based increased size net is compared to the current version. Fixed nodes in this case as system load would otherwise distort results. |
| Rosent | Winter | net_502_16x96 * | diff | 10.0+0.10 | LLR: -2.95 (-2.94, 2.94) [-5.00, 0.00] Games: 8948 W: 1999 L: 2170 D: 4779 Ptnml(0-2): 186, 1157, 1942, 1020, 169 | The larger 24x64 net architecture is failing STC so we instead try 16x96. This has far fewer parameters but still has a 5% slowdown we need to overcome. Based on training metrics this lies squarely between the baseline and 24x64. |
| Rosent | Winter | net_502_24x64 * | diff | 10.0+0.10 | LLR: -2.95 (-2.94, 2.94) [-5.00, 0.00] Games: 1566 W: 309 L: 453 D: 804 Ptnml(0-2): 65, 210, 333, 154, 21 | Having passed the fixed node test as expected, we try an STC test with [-5, 0] bounds, as the slowdown is expected to hurt more than in actual play. |
| Rosent | Winter | net_502_24x64 * | diff | N=40000 | LLR: 2.99 (-2.94, 2.94) [0.00, 5.00] Games: 2282 W: 745 L: 582 D: 955 Ptnml(0-2): 69, 209, 455, 306, 102 | Larger net, hopefully a clear improvement at fixed nodes. Around 16% less Nps. |
| Rosent | Winter | net_311m * | diff | 8.0+0.08 | LLR: -3.02 (-2.94, 2.94) [-2.50, 2.50] Games: 17246 W: 4156 L: 4292 D: 8798 Ptnml(0-2): 402, 2127, 3657, 2079, 358 | Trying to better understand the hybrid loss. Both nets are not fully trained, but are at a similar point in training. |
| Rosent | Winter | net_311l * | diff | 8.0+0.08 | LLR: -3.10 (-2.94, 2.94) [-3.00, 2.00] Games: 17742 W: 4218 L: 4374 D: 9150 Ptnml(0-2): 424, 2152, 3826, 2094, 375 | Trying to better understand the hybrid loss. Both nets are not fully trained, but are at a similar point in training. 311n removes the CE loss component completely. Even if worth Elo, we may not want this. |
| Rosent | Winter | net_311l * | diff | 8.0+0.08 | LLR: -2.97 (-2.94, 2.94) [0.00, 5.00] Games: 23074 W: 5479 L: 5505 D: 12090 Ptnml(0-2): 487, 2821, 4936, 2817, 476 | Final net with less CE loss component. |
| Rosent | Winter | net_311d * | diff | 8.0+0.08 | LLR: -2.95 (-2.94, 2.94) [0.00, 5.00] Games: 20226 W: 5005 L: 5044 D: 10177 Ptnml(0-2): 475, 2478, 4222, 2487, 451 | Reduced the weighting of the draw prediction loss component. |
| Rosent | Winter | net_311d * | diff | 8.0+0.08 | LLR: -2.95 (-2.94, 2.94) [0.00, 5.00] Games: 11346 W: 2730 L: 2809 D: 5807 Ptnml(0-2): 236, 1434, 2411, 1357, 235 | Draw calibration handled with MSE loss directly on the draw probability should be less sensitive to outliers that are common in the high entropy STC time controls the training data stems from. |
| Rosent | Winter | net_311 * | diff | 40.0+0.40 | LLR: 2.96 (-2.94, 2.94) [0.00, 5.00] Games: 3948 W: 860 L: 731 D: 2357 Ptnml(0-2): 32, 395, 1004, 498, 45 | Fully trained model with identical hyperparameters to prior gen, but extra data. This should clearly beat the master branch. |
| Rosent | Winter | net_311 * | diff | 8.0+0.08 | LLR: 2.96 (-2.94, 2.94) [0.00, 5.00] Games: 8158 W: 2101 L: 1925 D: 4132 Ptnml(0-2): 193, 922, 1694, 1056, 214 | Testing net with same hyperparameters but extra self play data a bit prematurely to get a sense of where things lie. |
| Rosent | Winter | net_test * | diff | 40.0+0.40 | LLR: 2.95 (-2.94, 2.94) [0.00, 5.00] Games: 3618 W: 807 L: 675 D: 2136 Ptnml(0-2): 37, 366, 880, 480, 46 | Trained with new data. The new training data has a very different distribution and has neutral effect on validation metrics. |
| Rosent | Winter | net_test * | diff | 8.0+0.08 | LLR: 2.97 (-2.94, 2.94) [-2.00, 3.00] Games: 7204 W: 1851 L: 1710 D: 3643 Ptnml(0-2): 160, 820, 1495, 973, 154 | Trained with new data. The new training data has a very different distribution and has neutral effect on validation metrics. |