totally fair! i spent a while trying to get clever and compress what i wanted to say and finally just hit submit but prob lost too much - local inference runtime, one cli that runs image/video/music/speech/3d/etc on your own machine without package hell. would def edit if I still could after your feedback, but window is closed. Thank you!
some real numbers on my m4 max -
image gen (zimage-nano), 1024x1024: 58s
image -> textured mesh (trellis.2): 2m 49s
sfx generate (5s clip): 3.6s
music generate (8s, ace-step): 15s
speech synth: 13s | transcribed back: 2.2s
video gen w/ audio(ltx unified-av) 4s,768x512: 2m 48s
text chat (laguna xs2.1): 102 tok/s - https://mlx.fast leaderboard
FYI, the title makes no sense in isolation.
totally fair! i spent a while trying to get clever and compress what i wanted to say and finally just hit submit but prob lost too much - local inference runtime, one cli that runs image/video/music/speech/3d/etc on your own machine without package hell. would def edit if I still could after your feedback, but window is closed. Thank you!
At least add “models”. Doing “music” on your local machine can mean many things.
Yes, missed that completely... Great point!