I have a habit of starting projects because I want one very specific thing.
Then somewhere along the way I apparently decide, "Well, if I'm alread...
For further actions, you may consider blocking this person and/or reporting abuse
okay yeah this is so cool😭
the whole lineage/take history thing is such a smart choice because AI workflows get messy SO fast. being able to look back and actually understand “wait, which prompt / stem / take got me here?” is huge.
also massive respect for normalizing all those completely different model interfaces into one workflow because that is exactly the kind of engineering problem nobody thinks about until they try to actually build something usable lol
and the “sometimes software can just do the normal thing” philosophy is honestly my favorite part. trimming, fading, exporting, etc. absolutely do not need AI sprinkled on top just because AI is involved somewhere else. it makes the whole project feel way more intentional instead of AI-for-the-sake-of-AI.
I'm looking forward to working with local models, though I have no clue where to start lol!
seriously cool project 💚
Ahh thank you 💖 and YES, the lineage/history thing became way more important than I expected once I started actually using it. AI audio gets messy fast when you’ve got generations, repaints, stems, converted vocals, mixes, etc. and suddenly you’re like “wait... what the hell did this come from?” 😂
And the different model interfaces were definitely one of those “this seems simple until you actually build it” problems. They all want basically the same musical ideas expressed in completely different ways, so making that feel like one coherent workflow was a huge part of the project.
Also I’m 100% with you on the normal-software part. I really didn’t want to shove AI into every single feature just because the app uses AI elsewhere. Sometimes trim should just trim lol.
And honestly local models are way less scary once you start playing with them. I have a feeling you’d have a ridiculous amount of fun with them 😂💖
Really interesting project. The part that stood out to me is not only the music generation, but the workflow around keeping track of prompts, takes, stems, and edits. Local AI tools become much more useful when they feel like real creative environments instead of just a button that generates output.
Also, separating the app layer from the inference backend is a great architectural choice. It makes the system more flexible and easier to evolve as models keep changing.
Thank you! That’s really what I’ve been trying to build toward. I didn’t want Miso to just be a prompt box with a Generate button, because once you start actually making things, the workflow around the output matters just as much as the generation itself.
Keeping prompts, takes, stems, edits, and where everything came from has ended up being a huge part of making it feel like a real creative tool instead of just a model demo.
And I’m really glad I separated the app from the inference backend early. The models are changing constantly, so having Miso depend on its own internal contracts instead of wiring the whole UI directly to one runtime makes it a lot easier to swap things around and keep the rest of the app stable.
Exactly. That distinction between a model demo and a creative tool is important. Once you keep prompts, takes, stems, edits, and provenance, the output stops being a disposable generation and becomes part of a workflow you can return to, compare, modify, and build on.
I think that’s where local AI gets really interesting. The model is only one component. The surrounding system determines how useful that capability becomes in actual creative work.
The lineage tracking is the part most people skip and it's the part that actually matters. I built something similar for image generation last year and the "which params produced the one I liked" problem ate more time than the generation itself — ended up with 600+ files in a flat directory before I gave up and wired up a database. 40 GB of models is brutal but honestly sounds about right for this scope.
A repaint can improve the selected section and still fail at its edit boundaries because the replacement may not share the original phase or ambience. Did you test those joins separately from the section's musical quality?
That’s a really good point. I tested the repaint results mostly for prompt responsiveness and whether the replacement itself behaved the way I expected, but I didn’t separately measure the edit boundaries for phase/ambience continuity.
In practice I’ve mostly been judging the joins by ear so far, especially on shorter repaint regions. That’s definitely something worth testing more deliberately though, because a musically good replacement can still sound wrong if the seam gives it away.
The "documentation says X but the model actually responds to Y" line hits home - ran into the same thing wiring local model backends into CogniRunner, where a parameter was silently ignored unless it was nested exactly where the model's own reference implementation expected it, not where the API schema claimed it belonged. The Docker split keeping playback and editing alive while the inference container is down is the right call too - a lot of local-AI tools tie the UI to the same process serving the model, and the whole thing locks up the moment a generation job backs up.
Yes, exactly. The silent-parameter thing is especially nasty because nothing necessarily fails. You get a valid result back, it just completely ignored the setting you thought you changed, which can waste a ridiculous amount of time.
And that separation between Miso and the inference backend ended up being one of the better architectural decisions in the project. I really wanted the library, playback, editing, exports, etc. to still feel like normal software even if the model server is busy or completely down.
Local model tooling is full of these tiny integration traps that only show up once you stop doing single happy-path calls and try to build an actual usable app around them 😅
Love it ! I always wonder that is there a app which can automatically mix the audio by just saying. Example may be a audio mix for Hip hop dance competition and iterate atop !
That would actually be really cool. Miso doesn’t do automatic mixing like that right now, but I love the idea of being able to describe the kind of mix you want and have it handle the transitions, pacing, levels, etc. A hip hop competition mix is a perfect example. You may have just given me another feature idea!
Awesome ! Appreciate you for taking inputs !
The "tiny experiment that got out of hand" pattern is one I recognize from my own projects. I started with a simple eval script and ended up with a full pipeline. The local-first approach is interesting for evals too: running offline means you can test without worrying about API rate limits or data privacy. What was the hardest part of making the local models good enough for real use?
Honestly, the hardest part wasn’t really making the models “good enough” so much as making them predictable enough to build around.
A lot of the work was figuring out what the models and runtime actually respond to versus what the docs imply. I hit cases where parameters were accepted but silently ignored, routes behaved differently than their names suggested, and memory usage was way higher than the model file size made it look. ACE-Step, for example, really needs its memory-saving mode on a 16 GB card, and some other families need their components loaded in a very particular way just to fit.
The other big part was normalizing all the little incompatibilities around them, like sample rates, prompt formats, output types, and task-specific request shapes, so the user doesn’t have to think about any of that. Miso ends up translating one workflow into whatever each model family actually expects.
So I’d say the hardest part was less “make the models better” and more “make the weirdness around the models disappear enough that the app feels reliable.”
the guided builder translating one set of controls into whatever syntax each model wants is the real product here. one model wants a caption, another wants tags, ACE-Step splits style and bpm and lyrics into separate fields for the same idea. that translation layer is invisible work nobody notices until it breaks and you're debugging why a param got silently ignored.
Exactly. That translation layer has probably been one of the most deceptively hard parts of the whole project.
From the user side it’s just “style, tempo, vocals, lyrics,” but under the hood those same ideas can mean completely different request shapes depending on the model. One wants a caption, another wants tags, another wants separate fields, and then some parameters are technically accepted in the wrong place and just get silently ignored, which is extra fun 😅
I really wanted the builder to hide the annoying syntax differences without flattening the models into pretending they all work the same way, because they definitely don’t.
So yeah, I agree. A lot of the real product is that invisible glue layer that keeps all the weird model-specific behavior from leaking into the workflow.
the repaint finding is the one I would have saved weeks on: surrounding audio influences the replacement more than the prompt does. we hit the same effect in AI music video pipelines where section swaps behaved completely differently depending on what flanked the cut point.
the prompt syntax divergence is the harder problem than it looks. ACEStep wanting BPM as a separate field while another expects comma tags is the same format negotiation problem you see in agentic tool calling. the abstraction layer is actually the hard part.
does the repaint hold across genre gaps, or does it break when surrounding sections are rhythmically far from the replacement prompt?
That’s pretty much where I’m at with it right now. Repaint works, but I wouldn’t say it holds genre changes especially well yet.
The surrounding audio seems to have a really strong pull on what the replacement becomes, so if the prompt is asking for something rhythmically or stylistically far away from what’s around the cut, the result can kind of get dragged back toward the original context.
I’m still figuring out exactly where that boundary is, though. That’s part of why I still think of Miso as very much a project in progress. There are a bunch of behaviors like this that only really become obvious once you use the models enough and start testing the weird edges.
And yeah, the abstraction layer has honestly been one of the harder parts. The models all want slightly different versions of the same musical idea, so making that feel coherent without hiding the important differences is a whole problem by itself.
I'm interested on getting more feedback from others.
The audio.cpp-on-a-separate-machine split is the detail worth stealing - browser talks to Miso, Miso talks to inference, so the app half doesn't care whether the GPU box is even up. I've been doing something close to that splitting compute across two Macs on a LAN, and decoupling the transport like that is what makes the rest (editing, export, playback) keep working when the model server doesn't.
On the repaint quirk - sounds like repaint starts denoising from something closer to the original waveform than a fresh noise draw would, so there's less room left for the prompt to move it. Did ACE-Step's code confirm that, or is it still inferred from the opposite-prompt test?
Yeah, the backend split has turned out to be one of those decisions that keeps paying for itself. Once playback, editing, export, project state, etc. don’t depend on the inference process being healthy, the whole app feels a lot less fragile.
On repaint, that’s still an inference from behavior/testing rather than something I confirmed directly in ACE-Step’s code. The opposite-prompt test was what really convinced me the surrounding audio is doing most of the steering there, but I haven’t traced the repaint implementation deeply enough to say exactly how the starting latent/noise state is constructed.
Your explanation would fit the behavior really well though, so now I kind of want to go dig through that path and see if that’s actually what it’s doing.
Sharded the model across two GPUs after wiping 12GB VRAM to avoid OOM; the audio tokenizer’s hallucination trap was the real nightmare. Any other local AI music hacks you swear by?