Hi,
Thanks for your work and the publishing of resources.
I would like to ask that how do we customize the text-to-image inference process, involving:
- specifying the output resolution.
- specifying the number of inference steps.
- turning off flash-attn when running.
Also, I find the performance of the Dif-DTok version impressive. Have you published the code for this diffusion detokenizer version?
Thank you in advance.
Hi,
Thanks for your work and the publishing of resources.
I would like to ask that how do we customize the text-to-image inference process, involving:
Also, I find the performance of the Dif-DTok version impressive. Have you published the code for this diffusion detokenizer version?
Thank you in advance.