DEV Community

Mike
Mike Subscriber

Posted on

✍️ MagenticCMS: Rehearse the Comments Before You Publish the Post

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange.

What I Built

Every social team has seen it happen: a post that looked clever in the draft gets screenshotted, quote-posted and turned into a backlash within hours. Sometimes it is a joke that punches down, a promotion that ignores a religious holiday, or a claim that sounds like a guarantee. The fix afterwards is expensive: deleting the post, writing an apology, and dealing with the damage to trust. Most of these mistakes are easy to spot if someone outside the team reads the draft first. The problem is that there is rarely time to ask.

MagenticCMS homepage: a place to write and rehearse a draft before publication

MagenticCMS is a pre-publication rehearsal room for content editors. Before a post goes out, you show it to a simulated audience and watch them react: who scrolls past, who laughs, who gets angry, who replies to whom, and which exact words set things off. An analyst then compares those reactions with the brand's own history (its red lines and past incidents, stored in Sanity), points to the phrase causing trouble, and suggests a rewrite. You can apply it, run the audience again, and compare.

The point is to catch avoidable mistakes early and make the post land better, while it is still a draft. It does not predict virality or replace real research: the readers are simulated personas, and the scores are rough signals for an editor to investigate, not forecasts. But an extra set of eyes before publishing is a lot cheaper than an apology after it.

A recorded simulation showing the 3D audience, reactions and conversation feed

What's inside:

  • A web composer: title, copy, an optional description of the image/video, platform, brand and test audience.
  • A live conversation feed: simulated readers scroll past, react, comment, and in later waves reply to each other. Reactions stream in with floating emoji and an optional beat, livestream-style.
  • A 3D audience table (Three.js): one figure per stored reaction, grouped by expressed sentiment, with a separate lane for people who scrolled past. Click a figure to read its comment and follow its real reply links.
  • An analyst: reads the reactions, the worst-reacting cohorts and the brand's own memory (red lines, past incidents stored in Sanity), quotes the exact phrase causing trouble, and drafts a rewrite.
  • An editorial workflow stored as data: agents may flag or clear a draft; only a human may approve or publish.

The audience is 379 persona records projected from a sample of the MatrAIx Persona-1M dataset (human-grounded and synthetic profiles). They are not real participants, not verified identities, and not a representative sample of any country.

The name nods to Microsoft's Magentic Marketplace, a multi-agent market simulation. MagenticCMS applies the idea to content, with the CMS behind it.

Demo

🔗 Live:

✍️ MagenticCMS ✍️

Sanity-powered content rehearsal: test drafts with AI personas, explore reactions and replies in 3D, and review suggested edits before publishing.
Live Demo 🔥 - https://magentic-cms.vercel.app

🎬 Video:

💻 Code: https://github.com/longphanquangminh/magentic-cms

There are only two controls: ▶ Run simulation (a new simulated audience; uses the LLM) and ↻ Replay (plays back a saved session; no LLM calls). Hype mode and sound are on by default and can be turned off with one click.

A quick tour that costs nothing:

  1. Open Saveo, 'Still broke at 30?' promo, select a completed session from its run history, and press Replay. You're watching stored results, not a new prediction.
  2. Drag the 3D table, click a figure, follow a reply link.
  3. Open Insights for the flagged phrases, the matching brand memory and the suggested rewrite.
  4. Open Workflow for the editorial history, with 🤖 and 🧑 entries side by side.

To test your own draft, use + Write a post (or open Weekend reading corner, invitation, a post I wrote in the composer), save it, then press Run simulation. Saving never calls the model. The demo is shared, so please don't enter private content. "Publish" here changes a CMS workflow state; nothing is sent to a social network.

Writing a new draft in the web composer, with a live feed preview

What it caught

The seeded brands and their incident histories are fictional test fixtures. One deliberately bad Saveo post read:

"Still broke at 30? 😂 Maybe skip the daily latte, bestie… No excuses."

Saveo's brand record contains a fictional 2025 "avocado-toast ad" incident and a red line against shaming people for small pleasures. The analyst connected the new wording to that stored history instead of producing a generic sentiment summary. In two 40-persona development runs the backlash index was 89/100 for the original and 42/100 after applying the suggested rewrite. Those are heuristic simulation indices: not percentages, not validated predictions, and not a controlled A/B test (a new audience was sampled for the second run).

For contrast, a separate Saveo draft took a more empathetic approach: "Prices went up. Your paycheck didn't… No judgment, just receipts." In a development run, it scored 18/100 and was cleared for review.

These examples explore different editorial choices, not validated predictions of audience behaviour. The useful output is the reasoning and the specific wording an editor can inspect, rather than the score alone.

Image d

When I tested an approval request labelled as coming from an agent, the workflow rejected it:

{"error": "\"Approve\" can only be taken by: human (not agent)"}
Enter fullscreen mode Exit fullscreen mode

Code

💻 Repository: https://github.com/longphanquangminh/magentic-cms

GitHub logo longphanquangminh / magentic-cms

Sanity-powered content rehearsal: test drafts with AI personas, explore reactions and replies in 3D, and review suggested edits before publishing.

MagenticCMS

V3: Web composer at /compose, interactive Three.js audience, and resumable API-failure handling. Updating an existing copy? Preserve your env files, run npm install in web, and restart. Do not reseed your existing dataset.

Rehearse the crowd before you post. Your draft goes "live" to a room of persona agents grounded in real survey records (MatrAIx Persona-1M). They react, comment and argue with each other in real time. An analyst agent reads your brand's memory (red lines, past incidents), flags the exact phrase that will get you dragged, and drafts a fix. A workflow stored as data in Sanity decides who may move the post forward: agents can flag and clear, only humans can approve.

Built for the DEV × Sanity Challenge (Path Two).

magenticcms/
├─ studio/     Sanity Studio: schemas, workflow actions, live "Simulation" tab, stage badges
├─ web/        Next.js live room + simulation engine (API
…
magenticcms/
├─ web/                         Next.js 16 app: live room, composer, simulation engine
│  ├─ app/
│  │  ├─ page.tsx               Home: posts, latest scores, dataset counts
│  │  ├─ compose/page.tsx       Write / edit a draft
│  │  ├─ posts/[id]/page.tsx    The live room
│  │  └─ api/
│  │     ├─ composer/           Create/update posts (optimistic _rev check)
│  │     ├─ posts/[id]/         Post state · workflow transitions
│  │     └─ runs/               Create run · step · finalize · cancel · reactions · SSE stream
│  ├─ components/
│  │  ├─ LiveRoom.tsx           Orchestrates a run from the browser, reveals reactions
│  │  ├─ AudienceWorld.tsx      Three.js audience table + 2D fallback
│  │  ├─ PostComposer.tsx       The editor
│  │  └─ useLiveAudio.ts        Web Audio: synthesized cues + 128 BPM beat
│  ├─ lib/
│  │  ├─ engine.ts              Sampling, crowd batches, analyst, transitions
│  │  ├─ workflowEngine.ts      Enforces the workflow document
│  │  ├─ metrics.ts             Pure scoring, shared by browser and server
│  │  ├─ modelTransport.ts      Bounded retries / failover, no raw provider errors
│  │  ├─ crowdValidation.ts     Rejects incomplete or duplicated model output
│  │  ├─ heal.ts                Closes runs abandoned by a closed tab
│  │  └─ sanity.ts · queries.ts · gemini.ts
│  └─ tests/resilience.cjs      Offline tests for the failure paths
├─ studio/                      Sanity Studio v6
│  ├─ schemaTypes/              post · brand · audience · persona · simulationRun · reaction · workflow · workflowEvent
│  ├─ actions/workflowActions.tsx   Document actions generated from the workflow document
│  ├─ components/SimulationPane.tsx Live "Simulation" tab (client.listen)
│  ├─ components/StageBadge.tsx
│  └─ structure.ts              Posts grouped by stage, runs, personas, workflow log
├─ scripts/
│  ├─ build_personas.py         MatrAIx parquet → persona documents (NDJSON)
│  └─ seed.py                   Personas, workflow, brands, audiences, demo posts → Content Lake
└─ DATA_ATTRIBUTION.md          Dataset terms and responsible-use notes
Enter fullscreen mode Exit fullscreen mode

My Build Process

MagenticCMS is my project: the idea, the product decisions, the testing, and the long list of things I sent back because they weren't right. I built it by directing an AI coding agent, running every version locally, and deciding what shipped. The prompts quoted below are paraphrases of my real requests, not a verbatim transcript.

Where it started

Before this challenge I had prototyped a "simulate a population before you ship" idea for another hackathon. Its weakness was obvious once I used it: it only told you things. For this build my rule was that the system must move content through a real editorial process: a person approves, the agent can only flag or clear. That's why the workflow lives in Sanity next to the content instead of in code.

I chose MatrAIx persona records as the audience source because I wanted reactions conditioned on structured attributes (income, politics, religiosity, attitudes to brands and influencers) rather than a generic "the internet thinks…" paragraph.

Data first: Python → Content Lake

scripts/build_personas.py reads the MatrAIx sample shard (999 rows × 994 columns) with pandas/pyarrow, keeps adults only, drops the Wikipedia-derived profiles of public figures, ranks records by how many of the attitude/personality fields are actually populated, and projects ~40 fields (Big Five/BFI-2 traits, attitudes to social media, influencers, brands, consumerism; income, region, generation, religiosity, shopping style, tone…) into a persona document. Each document keeps matraix{source, recordId, rowIndex} so every simulated comment can be traced to its source record. Handles are deterministic aliases generated from a seeded hash, not names from the data.

scripts/seed.py writes everything through the Sanity HTTP mutation API: 379 personas in batches of 100 (createOrReplace), the workflow document, two fictional brands with red lines and past incidents, four audiences, and three demo posts (createIfNotExists, with a --reset-posts flag to put them back to draft).

The schema is the product

post ──► brand            (voice, redLines[], pastIncidents[{title, date, whatHappened, lesson}])
  │  └► audience          (title, groqFilter over personas, sampleSize)
  │
  ├── simulationRun       (revision, bodySnapshot, mediaSnapshot, queue[{persona→, wave}],
  │     │                  status, processed, metrics{…}, cohorts[], analysis{mode, verdict,
  │     │                  headline, summary, riskFlags[{phrase, why, severity, brandMemory,
  │     │                  evidence[→reaction]}], suggestedRevision, changes[]})
  │     └── reaction ×N   (persona→, wave, seq, action, emoji, text, sentiment, likes,
  │                        replyTo→reaction, flag{category, reason}, thought)
  │
  └── workflowEvent ×N    (transition, from, to, actor{kind: human|agent, name}, note, run→)

persona                   (379 docs, read-only; traits{}, attitudes{}, matraix{source, recordId})
workflow.socialPost       (stages[], transitions[{id, from[], to, actors[], requiresNote}])
Enter fullscreen mode Exit fullscreen mode

Decisions behind it:

  • An audience is a GROQ filter, not a list of IDs. "Budget-conscious" is literally householdIncome in ["<$25k","$25k-50k"] || shoppingStyle == "Bargain hunter". The engine samples personas by running that filter against the Content Lake, so an editor can define a new audience in Studio without touching code.
  • replyTo is a reference from one reaction to another. That reference is what makes the agents talk to each other rather than beside each other, and it is exactly what the 3D view draws as reply links.
  • Every run snapshots the copy and the media description. Revisions stay comparable after the post changes.
  • Each risk flag carries evidence[] references to the reactions that justify it. The phrase is highlighted in the post; the comments behind it are highlighted in the feed.
  • Every persona points back to its MatrAIx record. Click "why?" on a comment to see the persona card, its private "thought", and the source.

Workflow as data, enforced in one place

Image Workflow

workflow.socialPost is a normal document you can open in Studio:

start_simulation   draft | needs_revision | ready_for_review → simulating   [human, agent]
flag_risk          simulating → needs_revision                              [agent]
clear_for_review   simulating → ready_for_review                            [agent]
apply_revision     needs_revision | ready_for_review → draft                [human, agent]
approve            ready_for_review → approved                              [human]
override_approve   needs_revision → approved                                [human, note required]
publish            approved → published                                     [human]
send_back          ready_for_review | approved | needs_revision → draft     [human]
Enter fullscreen mode Exit fullscreen mode

The web analyst and live-room controls use transition() in web/lib/workflowEngine.ts. Studio actions read the same workflow definition through their own client-side implementation. The web engine checks the current stage and actor label, patches the post and its draft when present, and appends a workflowEvent in one transaction. Studio actions also record the transition history.

Customising the Studio

Image xx

  • Document actions are generated from the workflow document: Run simulation, Apply suggested revision, Approve, Approve anyway (asks for a reason), Publish, Send back. Each one hides itself unless the current stage and a human actor allow it.
  • A "Simulation" view on every post: a custom pane that subscribes with client.listen() to the post's runs, so the risk/trend meters move while a run is in progress, even when it was started from the web app.
  • Stage badges and a structure tree that groups posts by stage with GROQ-filtered lists; separate lists for runs, reactions, personas and the workflow log; Vision for ad-hoc queries.
  • Personas, runs, reactions and the log are read-only and have no "create new" templates. Editors shouldn't hand-tune the crowd until it says what they want.

The engine

  1. Sample. Run the audience's GROQ filter, shuffle, cap at 120, split into three waves (50 / 30 / 20 %). The queue is stored on the simulationRun.
  2. Step. Each call to /api/runs/:id/step takes up to 16 personas from the same wave and plays them with Gemini (gemini-3.5-flash-lite, 8 personas per call, 2 calls in parallel, JSON mode with a response schema). Personas in waves 2 and 3 are shown the top and newest comments with IDs and may reply to or like them. Every model response is checked against the exact persona IDs requested; incomplete or duplicated batches are rejected. Accepted reactions, like increments and the queue update go into one transaction with an ifRevisionId lock on the run, so concurrent requests cannot commit duplicate results from the same queue revision. This protects stored evidence, not necessarily the cost of duplicate in-flight model requests.
  3. Stream. /api/runs/:id/stream subscribes to Sanity's listener on the server (the token never reaches the browser) and forwards "changed" pings; the browser pulls the delta by seq and reveals reactions one by one.
  4. Score. metrics.ts is a pure function used by both browser and server, so the live gauge equals the saved number. The backlash index weights negative sentiment, a scaled and capped flag rate, angry/sad reactions, and negative amplification (the share of comment likes that went to negative comments). The implementation clamps and rescales the result to 0–100. The trend signal combines engagement, share rate, reply depth and intensity. Cohorts are grouped by generation, region, political lean, income and religion (min n = 3).
  5. Analyse. gemini-3.5-flash gets the metrics, worst/best cohorts, the 14 most negative and 6 most positive comments with IDs, and the brand document. It returns a verdict, flags (verbatim phrase + brand memory + evidence IDs) and a rewrite.
  6. Move. If the analyst is unavailable, risk ≥ 45, any flag is high severity or the verdict is kill → flag_risk; otherwise clear_for_review.

For 40 personas that is normally 6 crowd calls + 1 analyst call. This is batched persona role-play, not one autonomous process per person, and the reveal animation is not token streaming.

Failure handling: "never invent evidence"

Asking "what happens when the API fails?" exposed a real bug: an early version counted a missing model response as a scroll, so an outage looked like audience behaviour. Now:

  • modelTransport.ts has bounded timeouts and retries, an optional explicit fallback model, and never surfaces raw provider error bodies.
  • Failed or incomplete batches stay pending; the run pauses with saved progress and offers Retry pending agents, Stop & return to draft, or Replay last completed session.
  • If only the analyst fails, the UI shows METRICS ONLY: deterministic scores, no invented rewrite or approval, human review required.
  • When Home or a post page is loaded, heal.ts checks for unfinished sessions with no progress for over 3 minutes. It closes them as interrupted, keeps their reactions, and returns the affected post to draft. This is load-triggered cleanup, not an unattended background scheduler. A fresh running session resumes on reload.
  • tests/resilience.cjs covers missing key, auth and quota errors, failover, invalid JSON, omitted/duplicated personas, partial-batch pause/resume and the metrics-only analyst path, all against an in-memory store with no real API calls.

Recovery UI under a simulated API outage: pending responses are not converted into invented reactions

The 3D room

AudienceWorld.tsx renders a tabletop with an orthographic camera and OrbitControls. One miniature (head + torso + base) per stored reaction, coral/stone/lime by sentiment, a separate lane for scrolls. Rendering is on demand (no idle loop), hover and click use raycasting, the selected figure's replyTo parent and children are drawn as curved tubes, and everything is disposed on unmount. A keyboard-accessible participant menu and a 2D cohort fallback keep the same data when WebGL is unavailable.

A recorded simulation showing the 3D audience, reactions and conversation feed 2

What I pushed back on

  • "It looks busy but means nothing." The first room was a dense dashboard; the next put decorative dots behind a floating post that overlapped the UI on long copy. I asked for 3D only if it carried data.
  • "Why can't I write my own post?" Early builds could only test seeded examples. The composer came from that, including an editable media description. Changing a claim in the copy does not fix a conflicting claim in the attached visual context.
  • "Three buttons is confusing." I cut the controls to Run simulation and Replay, and moved the rewrite loop into Insights.
  • "It should feel like a livestream." Hype mode speeds up the reveal of stored reactions, emoji float up per real reaction, and a synthesized beat (Web Audio, no files) plays from the moment a run starts until it finishes. None of it adds model calls or fake participants.

Bugs I found while testing

  • Emoji vanished on my machine: Windows had animation effects off and the CSS hid them under prefers-reduced-motion.
  • Posts stuck on "simulating" after F5: fixed with resume-on-reload and the 3-minute interrupted-run cleanup.
  • The beat stopped the moment a simulation started: resetting the feed silenced it and nothing restarted it; now one rule decides when audio plays and I tested start, analysis, replay, pause, sound-off and error paths.
  • Sanity specifics: listener filters can't follow references (run->post._ref), so the Studio pane listens to simulationRun documents; @sanity/icons v5 needs per-icon imports.
  • Smaller things: an off-centre logo glyph, an unreadable dropdown placeholder, a missing favicon, demo content that wasn't in English.

What I didn't do

  • I modelled the workflow myself as documents rather than using the new Workflows beta, which shipped two weeks before the deadline. Moving workflow.socialPost onto it is the obvious next step.
  • No App SDK app; the live room uses @sanity/client. No Context MCP or Knowledge Base; this is a Path Two entry.
  • The agents can't see images; they read the media description.
  • The browser drives processing steps; durable background execution is next.

Limits I'm aware of

  • Scores are heuristic indices, not probabilities, and are not validated against real feedback.
  • Persona attributes and LLM outputs can be biased. "Human-grounded" is not "verified". Cohorts overlap, and n = 3 says nothing about real groups.
  • The demo is shared and unauthenticated. Human/agent labels demonstrate workflow logic, not a production authorization boundary. Real authentication and quotas are required before any production use.

Sanity Project Details

  • Project ID: bh5o6glp
  • Dataset: production (public reads; server-side writes use a private token)
  • Document types: post, brand, audience, persona, simulationRun, reaction, workflow, workflowEvent
  • Query it yourself: https://bh5o6glp.api.sanity.io/v2025-02-19/data/query/production?query=*[_id=="workflow.socialPost"][0].transitions

The project ID is shared for judging the schema; it is not a credential. No write tokens or LLM keys are published.

Data attribution: MatrAIx Persona-1M · MatrAIx paper. The dataset is released for non-commercial research use only, including subsets and derivatives; upstream source terms apply. This is a research prototype, not a claim of commercial data rights.

Thank You, and What Could Come Next

A big thank you to @thepracticaldev and the Sanity team for running this challenge. MagenticCMS started as an idea I had been carrying around for a while, and this challenge gave me the push, and the deadline, to actually build it and find out where it breaks.

While building it, I kept thinking about who would use something like this. Anyone who publishes to a large audience, often without time to ask for a second opinion first:

  • Creators and public figures testing a post before it reaches their followers
  • Social media managers checking brand campaigns and announcements
  • Community admins and event organizers preparing updates, rules or calls for participation
  • Large publishing platforms like @thepracticaldev, where many authors share ideas with a global community

A rehearsal step could help them check a planned post, spot wording that might be misread, and try a stronger version before it goes live. The goal is not to guarantee a hit, but to learn something useful before publishing.

I can also imagine this as a small feature inside a blogging platform: a quiet "rehearse before publishing" button next to Save draft. Authors could preview how different readers might respond, notice unclear or risky phrasing, and still make the final call themselves. If the DEV team ever explores something like this for the editor, I'd love to help think it through, with clear consent, careful limits, and respect for privacy.

Thanks again for the opportunity to experiment. I'd love to hear what you think, especially from people who publish to large audiences.

Top comments (2)

Collapse
 
tri_nguyn_f5d8409891059 profile image
Dan Lopez •

woah mind blowded 🤯

Picked as gem
Collapse
 
kartik-nvjk profile image
Kartik N V J K •

Counting a missing model response as a scroll made an API outage look like audience behaviour. The 89 to 42 drop mixes the rewrite with a fresh audience, so I'd reuse the same personas for both drafts, which is how I keep my prompt A/B tests honest. Have you measured the index spread across reruns of one draft?