VQ-GAN model trained on the Kimetsu no Yaiba dataset on Tensorflow.
1model:
2 vqvae_config:
3 beta: 0.25
4 num_embeddings: 50257
5 embedding_dim: 128
6 autoencoder_config:
7 z_channels: 512
8 channels: 32
9 channels_multiplier:
10 - 2
11 - 4
12 - 8
13 - 8
14 num_res_blocks: 1
15 attention_resolution:
16 - 16
17 resolution: 128
18 dropout: 0.0
19 discriminator_config:
20 num_layers: 3
21 filters: 64
22
23 loss_config:
24 discriminator:
25 loss: "hinge"
26 factor: 1.0
27 iter_start: 50000000
28 weight: 0.8
29 vqvae:
30 codebook_weight: 1.0
31 perceptual_weight: 4.0
32 perceptual_loss: "vgg19" # "vgg16", "vgg19", "style"
33
34trainer:
35 batch_size: 64
36 n_epochs: 10000
37 gen_lr: 3e-5
38 disc_lr: 5e-5
39 gen_beta_1: 0.5
40 gen_beta_2: 0.9
41 disc_beta_1: 0.5
42 disc_beta_2: 0.9
43 gen_clip_norm: 1.0
44 disc_clip_norm: 1.0
Implementation and documentation can be found
here