This is reuploaded here for easier fuctioning with colab etc.
The following is a description made by eihpirtsa on their model page.
This Adetailer model will segment speech bubbles, text and watermarks commonly found in training data. Trained this so I could eventually automatically clean images in a dataset. Only tested on Comfy, but should work on other webUIs too. This is a WIP, and I have many things in mind on which could be improved:
-
make sure you don't set minimum confidence too low, or else undesired objects will be segmented
-
can misidentify watermarks for text, speech bubbles for logos etc. but this should not matter since they are segmented anyway
-
Some text that is transparent/partially hidden won't be identified
-
Trained primarily on NSFW images, may not work too well with comics, images with large/strange fonts etc.