Week 4 Dev in Progress

•

TL; DR

  • V1 cut down to seven steps, ending on publish or re-shoot? for playtest.
  • The full month (four in-game weeks) is finalized ~40 minutes, day by day.
  • Lip sync in generated video is fixed.
  • The MV pipeline ran successfully end to end.

· · ·

Cutting V1 to one clean run

Quarters gave us a build worth showing and a clear read on what wasn’t going to be finished. So we scoped V1 down to a single uninterrupted run.

What moved out:

CutWhere it went
F9 Expression classV2
F4 Dorm + F5 QuestsFolded into the Home scene
Everything after the video: publish growth, data report, ledger panelLater

· · ·

The 1 month in-game storyline:

You’re signed, but you haven’t debuted, and you have a month to earn it. Two things decide whether you do: your weekly evaluation grades, and how fast your following grows.

How it endsWhat it is
Week 1~13 minBeing found, chosen, seenV1 · the playtest build
Week 2~7.5 minFrom arranged-for-you to becoming someoneNext · music generation
Week 3~8 minExplore on your own; meet other players
Week 4~11.5 minYou stand on the evaluation stage and perform

Week 2 is the one we build next, and it’s written down to the day:

DayWhat happens
MonWriting your songYou brief the agency’s producer. It goes into production that takes a few days
TueVocal trainingPitch-matching minigame. Your own voice, or a voice chosen to represent you
WedThe song arrivesTwo demos, in a generic reference vocal. You pick one. Only then does the backend render it in your voice
ThuDance trainingA rhythm game charted from your own song
FriExpression trainingWeek 1’s memory game, with one difficulty step up
SatPostingPick one practice photo from the week, write something, post it
SunWeekly reviewThe producer’s report: this week’s grade and feedback

· · ·

Lip sync had been drifting: 1.2 seconds of silence

Original assumption: the video model being imprecise.

It wasn’t.

Seedance does not track the audio you hand it. It listens to the audio, re-speaks what it heard, and animates its own rendition of that. So it was never syncing to the player’s recording. It was syncing to its own interpretation of it.

And what made it mishear: roughly 1.2 seconds of phone room tone before the player’s first word.

The fix is word-boundary trimming through ElevenLabs Scribe: cut to ±150ms around the actual speech.

· · ·

The MV pipeline

We ran the music-video pipeline end to end for the first time:

Pick a real K-pop MV and measure it. Write the concept. Generate the song in Suno. Build the storyboard on the song’s actual bars. Generate the character look and turnaround, then the set plates. Then generate the video, segment by segment.

This is not only our backstage process. It’s the flow the player walks through in the game. Every step maps to a choice made in a UI.

· · ·

Cost Breakdown

ToolUsed forTypeCost
OpenAIArchitecture, codingAPI$13.33
OpenRouterConcurrency probes ~$8; lip-sync probe $0.76; F7 collage probe ~$0.30API$11.35
Pollo.aiVideo generationSubscription$129
SunoMusic generationSubscription$30
ElevenLabsWord-boundary trimming (ASR)Subscription$11.66
Claude: $125*6Team efficiencySubscription–
Total$195.34

Not approved cost: fal.ai ($2.93)

· · ·

Weekly Progress Report