Curated Formatv0.4.0

Animate Shaz

Give Shaz a voice track and pick a room. The kit reads the words locally, lip-syncs the mouth, and gives a fresh agent five artist-reviewed gestures for the moments that matter.

By Shaz · Updated August 2026

Before you start

No subscriptions. No API keys. It runs on Apple silicon.

Everything stays local

The rig, transcript, lip-sync, and video tools are all included in the download. Your audio never has to leave your machine. Node, FFmpeg, and Apple’s command-line tools are required.

$0 service fees

The workflow makes no network calls and uses no paid generation service.

It hears the words, too

Whisper 1.9.2 writes the transcript and word timing. Cherry 0.1.0 picks the mouth shapes. The same Shaz rig draws the scene.

Included assets

Five reviewed gestures. Local transcription and lip-sync. Four rooms.

The five gestures shown below were recreated from artist animation and reviewed as ready to use. The kit contains 14 runnable actions in all. One is the calm body behind Talk to Camera; the remaining 8 are engineering reference material and need a fresh creative review before a finished video uses them. The kit reads the spoken words locally, then Cherry maps the audio to five hand-drawn mouth shapes.

Understands the dialogue

Local English transcript

Whisper 1.9.2 writes the words and their timing on your Mac, so gestures can land on what Shaz is actually saying. No upload or API key.

Lip-sync included

Cherry Lip Sync 0.1.0

Cherry listens to the audio and chooses the matching mouth shape on your Mac. It works without a subscription, network call, or second animation system.

Four built-in backgrounds

Pick the room. Keep the camera fixed.

Sisters Room remains the main default. Living Room adds a warmer home setting, Photo Zone removes the old map artwork cleanly, and Pure White gives Shaz a neutral stage. The camera stays fixed in every room.

Sisters Room built-in Shaz background

Sisters Room

Default

main/default environment for Shaz dialogue and body-language video; fixed camera only

Living Room built-in Shaz background

Living Room

warm home environment for dialogue and body-language video; fixed camera only

Photo Zone built-in Shaz background

Photo Zone

Future media zone reserved

clean purple room with the original map artwork removed; the cleared area is reserved for future supporting media but is a fixed background in this release

Not active yet: no overlay, crop, replacement, or supporting-media input is implemented.

Pure White built-in Shaz background

Pure White

neutral pure-white environment for minimal scenes and downstream compositing; fixed camera only

Default dialogue option

Talk to Camera

Shaz faces the audience in a calm, grounded pose while the supplied audio changes only the mouth drawing. Use it for ordinary speech, then add Present, Think, Ah-ha, Point, or Confident only when the line earns a gesture.

InputsequencePreset: "talk-to-camera"

Keeps the body steady

neutral-listening

  • The audio decides the length
  • No manual frame math
  • The same body and renderer stay in place

The talking kit

Five hand-drawn mouths. Every sound has somewhere to go.

Cherry listens to the audio; Shaz swaps between five mouth drawings while the body, hands, timing, and room stay untouched.

Rest / closed hand-drawn Shaz mouth shape

01 · Cherry A · X

Rest / closed

silence + closed-lip consonants

Teeth / EE hand-drawn Shaz mouth shape

02 · Cherry B · G · I · J

Teeth / EE

teeth, EE, F/V + CH/J/SH

Small open hand-drawn Shaz mouth shape

03 · Cherry C · H

Small open

EH + tongue-forward L

Wide open hand-drawn Shaz mouth shape

04 · Cherry D

Wide open

wide AH

Rounded O hand-drawn Shaz mouth shape

05 · Cherry E · F · K

Rounded O

OH, OO/W + R

Only the mouth changes · the body stays put

Five artist-reviewed gestures

Shaz performing the reviewed Present, Think, Ah-ha, Point, and Confident gestures

01 · Artist-reviewed

present

02 · Artist-reviewed

think

03 · Artist-reviewed

aha

04 · Artist-reviewed

point

05 · Artist-reviewed

confident

Small supporting drawings

phone: look-at-phone · crossed-arms-pose: arms-crossed-skeptical registered pose drawing

Examples

The first video proved Shaz could talk and gesture. In the new 30-second story, a fresh agent uses the transcript to land Think on “idea,” Point on “least,” and Confident on “best.” Play both with sound.

How a talking scene gets made

Choose the performance → Check the plan → Render Shaz → Watch the result → Deliver

1. Choose the performance

Free

The kit reads the words and their timing locally. Start with Talk to Camera, then place an artist-reviewed gesture where the line needs more expression.

2. Check the plan

Free

Makes sure the audio, background, gesture IDs, timing, and rig files are ready before rendering.

3. Render Shaz

Free

Shaz’s original rig draws every body frame and mouth shape.

4. Catch visual problems

Free

Checks for broken joints, clipping, misplaced layers and props, facial glitches, wrong duration, and bad video settings.

5. You approve it

Free

A person watches the exact MP4 before the agent can deliver it.

Waits for you

Commands the agent runs

npm run check
npm run inspect:registry
npm run smoke
npm run transcribe -- --audio=/absolute/path/audio.wav --output=/absolute/path/transcript.json
npm run lipsync -- --audio=/absolute/path/audio.wav --output=/absolute/path/cherry.tsv
npm run init -- --run=episode-01 --input=/absolute/path/input.json --audio=/absolute/path/dialogue.wav
npm run validate -- --run=episode-01
npm run render -- --run=episode-01
npm run inspect -- --run=episode-01
npm run finalize -- --run=episode-01

02 · First-draft talking scene

A real voice track, performed by Shaz.

Open proof report
00:00 / 00:12
Open finished ad

03 · What gets checked

The agent checks the render. You judge the performance.

Open quality.json
0.4.0kit version
14runnable actions
5artist-reviewed gestures
$0service fees

A strong first draft, not a finished performance. The 12-second video above was made with an earlier 0.2.0 kit. It generated 100 mouth-timing cues locally, used 5 hand-drawn mouth shapes, and passed the audio, video, and rig checks in a fresh download. The current kit can also read English dialogue with word timing before it plans gestures, and includes Talk to Camera plus four built-in backgrounds. Creative review of this exact video is still pending.

Checks before review20 automatic checks, then 8 questions for a person
  1. 01source Xstage provenanceRequired
  2. 02compiled asset checksumsRequired
  3. 03pose recipe registry checksumRequired
  4. 04native-frame shoulder-to-sleeve-to-hand topology, or exact registered part-replacement tuple, component, exclusivity, and placement contract on replacement framesRequired
  5. 05native-frame hand-to-sleeve cuff ownership and contactRequired
  6. 06native-frame hand-to-sleeve and hand-to-head proportion limitsRequired
  7. 07per-frame clippingRequired
  8. 08per-frame finished layer order, including recipe-declared native arm crossovers, with hidden construction-arm and stray back-bang rejectionRequired
  9. 09per-frame facial pop limitsRequired
  10. 10pose-specific consecutive-identical-frame limitRequired
  11. 11exact prop presenceRequired
  12. 121280x720 H.264 yuv420p at 24 fpsRequired
  13. 13exact timeline durationRequired
  14. 14AAC presence and audible volume for audio-backed runsRequired
  15. 15registered background id, exact background checksum, and fixed camera for audio-backed runsRequired
  16. 16waist-up hoodie continuation below the bottom frame edge and clear horizontal margins across every used actionRequired
  17. 17local transcript checksum, source-audio checksum, canonical PCM checksum, ordered word timings, whisper.cpp source and build-plan checksums, locally compiled binary checksum, model checksum, and zero-provider receiptRequired
  18. 18gesture planning transcript SHA plus an exact transcript word and 24 fps frame anchor for every expressive audio-backed sequence entryRequired
  19. 19lip-sync cue source, audio checksum, cue checksum, bundled engine artifact checksum when locally generated, engine version, five-mouth mapping, and final resting mouth when lip-sync is requestedRequired
  20. 20talk-to-camera preset provenance, exact measured-audio frame count, exclusive neutral-listening body ownership, zero gaps, and contact-sheet samples at real mouth changesRequired

Every automatic check must pass. Then a person must watch the exact MP4 before final delivery.

04 · Everything included

Repo files

Open any file to read its actual contents.

Download exact Repo
Format manifestformat.json
{
  "id": "shaz-puppet-runtime",
  "version": "0.4.0",
  "title": "Animate Shaz",
  "summary": "Give Shaz a voice track, read the local word-timed transcript, choose one of four built-in backgrounds, and render a talking scene. Use Talk to Camera for everyday speech or add an approved gesture when the words earn one.",
  "mediaType": "video",
  "services": [],
  "providerCost": "$0",
  "officialRuntime": "runner.mjs",
  "sourceRig": {
    "kind": "Toon Boom Xstage reconstruction",
    "manifest": "rig-v2/runtime.json",
    "artistRenderedFramesUsed": false
  }
}
Agent instructionsSKILL.md
---
name: shaz-puppet-runtime
description: Animate the supplied Shaz puppet locally. Use Talk to Camera for dialogue, arrange approved body-language gestures, or repair and review one rig action without rebuilding the renderer.
---

# Animate Shaz

Skill version: **2.0**.

Use this kit to turn a voice track into a Shaz talking scene or to build a short performance from the recovered rig. Everything runs locally, makes no provider calls, and costs $0.

## Choose the job

- **Talk to Camera:** the normal choice for direct-to-audience speech. `sequencePreset: "talk-to-camera"` measures the audio, holds `neutral-listening` for the full line, and lets Cherry change only the mouth. Do not invent a pose or calculate frames.
- **Reviewed gesture sequence:** arrange the five artist-reviewed gestures listed below, then follow the complete run workflow.
- **Action repair or authoring:** work on exactly one action. Read `references/rig-animation-playbook.md` completely and follow the author-and-learn loop. Do not repair several unapproved actions at once.

## Which actions may be used

Use `neutral-listening` as the calm body behind Talk to Camera:

- `neutral-listening`

For body-language beats, default to these five artist-reviewed gestures:

- `present`
- `think`
- `aha`
- `point`
- `confident`

The registry also contains `shrug`, `key-point`, `excited-celebration`, `point-at-screen`, `look-at-phone`, `facepalm-frustrated`, `arms-crossed-skeptical`, and `phone-use-sequence`. They are runnable engineering material, not approved performance choices. Do not select one automatically or put it into a user video until that exact current recipe has passed a fresh complete visual review.

**Registered means runnable. It does not mean creatively approved.** Mechanical inspection can pass a pose that still looks wrong.

## Run workflow

1. Read `README.md`, `input-contract.json`, `composition-contract.json`, `output-contract.json`, `quality.json`, and `content-boundary.json`. Read `ROADMAP.md` as well when changing the Format, planning a capability, or checking what remains unfinished.
2. Run `npm install` once. Then run `npm run check`, `npm run inspect:registry`, and `npm run smoke`.
3. If the job uses audio, transcribe it before choosing gestures:

   `npm run transcribe -- --audio=/absolute/path/audio --output=/absolute/path/transcript.json`

   Read the full text and the word timestamps. The bundled English Whisper model runs locally; never upload the audio to Deepgram or another transcription service. For a gesture sequence, copy the printed transcript SHA into `planningTranscriptSha256`. Anchor each expressive action to a real `wordId`, label, and 24 fps frame (`round(startMs × 24 / 1000)`), then make the preceding frames add up to that anchor. Do not infer semantics from volume alone. Talk to Camera does not need a gesture plan, but initialization still records the local transcript as evidence.

4. Choose the input:
   - For ordinary dialogue, copy `fixtures/talk-to-camera/input.json`. Supply no `sequence`, `durationFrames`, or frame math. Initialization derives one exact-length `neutral-listening` hold from the audio, with lip-sync required.
   - For body language, write a `sequence` with the five reviewed gesture IDs above. Use `neutral-listening` only as the calm default or connective tissue. Use explicit `holdFrames` and `gapFrames`. The last action must use `gapFrames: 0`.
   - Choose `sisters-room`, `living-room`, `map-photo-zone`, or `pure-white` from `assets.json`. Name `backgroundId` explicitly for an audio-backed sequence. A semantic performance input may omit it and use `assets.defaultBackgroundId`. Never invent a background ID.
   - `map-photo-zone` is only a clean fixed room in this release. Its empty area is reserved for future supporting media; do not add, crop, or position an image or video there.
5. Start a run with an absolute input path:

   `npm run init -- --run=my-run --input=/absolute/path/input.json`

   For an audio-backed `shaz-sequence-input-v1`, also pass `--audio=/absolute/path/audio`. The bundled Cherry Lip Sync 0.1.0 WASI engine creates and records cues by default. Use `--lipsync-cues=/absolute/path/cherry.tsv` only for a real Cherry TSV generated from that exact audio. Use `--lipsync=off` only when the user explicitly wants audio without mouth animation. Never make cyclic placeholder cues or reuse cues from a different audio file.

   `shaz-body-language-performance-v1` remains body-language-only. Its audio sets duration and gesture timing, not mouth drawings. It accepts the same backgrounds and defaults to Sisters Room.

6. Run these commands in order:

   - `npm run validate -- --run=my-run`
   - `npm run render -- --run=my-run`
   - `npm run inspect -- --run=my-run`

7. Watch `agent-runs/my-run/final.mp4` completely. Inspect `contact-sheet.jpg` and `quality-report.json`. Do not approve a video you did not watch.
8. Edit only `agent-runs/my-run/human-review.json`. Keep the exact `reviewedOutputSha256`, name the reviewer, write concise notes, and set `status` to `approved` or `rejected`.
9. Run `npm run finalize -- --run=my-run`. Delivery remains blocked unless validation, inspection, file checks, and human review all pass.

## Author-and-learn loop

1. Choose exactly one action. Record its reference segment, meaning, duration, and visible acceptance criteria. Check the surrounding frames before choosing the cut: include setup through release, or give a smaller gesture a smaller, honest name. A preexisting filename does not define the action.
2. Watch the complete reference at normal speed. Then inspect dense decoded frames and consecutive-frame differences. When source controls exist, inspect their timing, hierarchy, drawing substitutions, deformation channels, and asymmetry. Find intentional stepped exposures and holds before deciding how to interpolate. Never infer motion from one destination frame.
3. Render the authored calibration through the official runtime before changing its meaning. Fix renderer-wide silhouette, deformation, masking, fill, and paint-order defects before tuning timing.
4. Animate through real rig controls and existing drawing substitutions. Preserve secondary controls and living holds. Make the smallest semantic change that produces the intended action.
5. Run focused tests and independent per-frame pose inspection. Where a reference exists, compare synchronized full-frame and close-up playback.
6. Watch the exact candidate completely at normal speed. Slow motion is useful for diagnosis, but it cannot be the only approval view. A successful command, contact sheet, or handful of frames is not visual approval.
7. Allow at most three candidate attempts. Fix only observed causes. Stop and report the blocker instead of weakening a gate.
8. After genuine approval, ask exactly: **“What did this teach us, and does the skill, runtime, or test suite need updating?”**
9. Record the behavior, root cause, smallest reusable correction, and evidence. Follow the promotion rules in `references/rig-animation-playbook.md`. Update a recipe checksum only after the accepted file is final.
10. Register and sequence the action only after inspection and checksum-bound human approval pass. Registration alone does not grant creative approval.

## Rules that keep Shaz on-model

- Protect the finished silhouette and character assembly before polishing motion. Visible seams, detached joints, missing fills, or construction artwork make timing judgments invalid.
- Reproduce the artist's timing grammar: anticipation, accent, overshoot, settle, readable living hold, afterbeat, and release. Do not apply generic smoothing.
- Preserve stepped presentation. If the reference animates on twos or holds an exposure, encode those steps instead of inventing smooth in-betweens.
- Preserve the whole control choreography. Major arm and head keys alone are rarely enough.
- Treat drawing substitutions, visibility, AutoPatch-style masking, and paint order as animation controls.
- Visible alpha and semantic ownership are not the same thing. A partial eye, hand, or face drawing may own a larger matte than its painted pixels. Rebuild that envelope and clip occluders behind it instead of shifting artwork or erasing only an outline.
- Palette ownership matters. Shape, opacity, and connectivity can pass while teeth, eye whites, tongues, or skin use the wrong color. Keep direct color-presence gates for stable semantic regions.
- Prefer limbs animated through their common rig ancestor, with one continuous native shoulder-to-finished-sleeve-to-hand chain per side. Independent screen-space pieces and multi-fragment limb assemblies are forbidden.
- If three bounded native-rig attempts prove that the recovered drawings and pivots cannot form an essential destination, one coherent part-specific drawing may replace the complete corresponding native parts. It must be user-supplied or explicitly authored for the pose, exact-file and exact-transform locked, placed on a declared paint layer, preserve the rig-rendered head and body, and never overlap visible native counterparts. This is a narrow substitution rule, not permission to use a full-character sprite.
- A front overlay hand may record its recovered native sleeve owner as descriptive, validated metadata, but that label does not create ownership. The renderer must derive the cuff owner from rig topology, matte the overlay at that finished cuff, and enforce the rendered hand role's geometry limits. Inspect the wrist joint; overlap with the face or torso is not attachment evidence.
- When reusing a gesture, keep its authored wrist states with its hand drawings. Do not import wrist transforms from another frame range or enlarge a hand to force contact.
- Establish handheld props before contact and keep the hand inside the native rig hierarchy. A prop cannot stand in for a hand, finger, fist, sleeve, or arm. Matching coordinates do not create a joint. A prop-free alias removes the prop while leaving the native gesture controls and drawings unchanged. Never use the part-replacement rule as a shortcut for prop interaction.
- Reject a candidate when dense transition frames or normal-speed playback show a detached, duplicated, missing, scale-popping, undersized, oversized, or independently drifting limb. Whole-character connectivity and an asset ID are not enough. Verify native shoulder, cuff, wrist, and proportions directly. For a registered replacement, also verify exact asset bytes, transform, paint layer, complete suppression of the native parts, and preservation of the original body and head.
- Diagnose in this order: assembly, deformation, substitution or expression, timing, then polish.
- Turn repeated mechanical failures into tests or inspection gates. Do not rely on an agent remembering prose.
- Use Talk to Camera for unaccented speech. Keep Present, Think, Ah-ha, Point, and Confident as brief meaningful accents, not constant motion.
- Keep the background separate from pose and mouth decisions. Every built-in room is fixed input data behind the same waist-up character path. Never add camera movement or a background-specific character transform.

## Hard boundaries

- `runtime/rig-v2-renderer.mjs#renderRigFrame` is the only renderer for smoke, proof, and final video.
- Lip-sync may change only the registered `Mouth` drawing at each output frame. It must not alter body controls, pose order, hold length, deformation, props, framing, or any other face drawing. Cherry runs locally through Node WASI as a cue generator, never as another renderer. Validation binds the exact engine artifact, cue file, and audio checksums.
- Do not download, invoke, or require a native `cherrylipsync` executable. Use `npm run lipsync -- --audio=/absolute/path/audio --output=/absolute/path/cues.tsv`, or let audio-backed sequence initialization call the same bundled engine.
- Transcription must use the bundled whisper.cpp source and English model through `npm run transcribe`. Never use Deepgram, another hosted service, an unsigned downloaded Whisper executable, or a source-supplied transcript as generated evidence. The transcript may guide pose choice and timing; it must never mutate the rig or renderer directly.
- The waist-up crop is part of the composition contract on every background. The hoodie must continue below the bottom edge through every used action, while hands and pointing gestures retain clear horizontal margins. Never reveal the rounded lower hoodie boundary or invent legs.
- Finished artist-rendered animation frames may guide phase, cadence, and acceptance criteria. They must never become runtime sprites, deformation data, or generated pose artwork. A user-supplied pose-design drawing may be registered only under the strict part-substitution rule above and must be disclosed separately from recovered Xstage drawings.
- Never bypass `poses/index.json` with an arbitrary recipe path.
- Never weaken per-frame clipping, joint continuity, layer order, prop, facial-pop, or provenance gates to force a pass.
- A run gets at most three render attempts. Fix the input or recipe between attempts; do not create shadow runs to evade the limit.
- Do not use a generic crossfade, whole-character bounce, or uniform interpolation in place of body mechanics.
- Do not copy only the obvious controls, approve from sparse screenshots, or patch several poses at once.
- A new action is an authoring task, not a sequence-input change. Follow `references/rig-animation-playbook.md` and `poses/README.md`, inspect the recipe independently, register its checksum, and add evidence before using it.

## Return the result

Give the user the absolute path to `final.mp4`, the video checksum from `delivery.json`, the selected background ID, the actions in order, `agent-runs/<run>/transcript.json` plus `delivery.json.transcript.sha256`, the anchor phrases used for body-language beats, and any remaining limitations. Never claim delivery when `delivery.json` is absent.
Runtime requirementsrequirements.json
{
  "schemaVersion": 1,
  "localTools": [
    { "name": "node", "minimum": "20", "required": true, "requiredFor": "official runtime, bundled Cherry WASI cue engine, and transcript workflow" },
    { "name": "npm", "required": true },
    { "name": "ffmpeg", "required": true },
    { "name": "ffprobe", "required": true },
    { "name": "xcrun", "required": true, "platform": "darwin-arm64", "requiredFor": "locate Apple Clang and the macOS SDK for the first local whisper.cpp build" },
    { "name": "clang", "required": true, "platform": "darwin-arm64", "source": "Xcode Command Line Tools", "requiredFor": "compile the bundled whisper.cpp source locally without a downloaded executable" },
    { "name": "tar", "required": true, "platform": "darwin-arm64", "requiredFor": "extract the bundled whisper.cpp source archive into the local build cache" },
    { "name": "zip", "requiredFor": "build:kit" }
  ],
  "bundledEngines": [
    {
      "name": "cherry-lip-sync",
      "version": "0.1.0",
      "artifact": "WebAssembly/WASI module",
      "host": "node",
      "nativeExecutable": false,
      "networkRequired": false,
      "purpose": "generate A-K/X speech cues for audio-backed shaz-sequence-input-v1 runs"
    },
    {
      "name": "whisper.cpp",
      "version": "1.9.2",
      "artifact": "checksum-pinned source archive plus base.en Q5_1 model",
      "host": "locally compiled Apple Silicon helper using Apple Clang and Accelerate",
      "nativeExecutableIncluded": false,
      "nativeExecutableBuiltLocally": true,
      "networkRequired": false,
      "supportedPlatform": "darwin-arm64",
      "purpose": "create an English transcript with word timestamps before body-language planning"
    }
  ],
  "environmentVariables": [],
  "providerCalls": [],
  "networkRequired": false,
  "networkRequiredForOperation": false,
  "networkMayBeRequiredForNpmInstall": true,
  "estimatedCost": "$0"
}
Input contractinput-contract.json
{
  "schemaVersion": "shaz-sequence-input-v1",
  "required": ["title"],
  "oneOf": [
    { "required": ["sequence"] },
    { "required": ["sequencePreset", "backgroundId"] }
  ],
  "properties": {
    "title": { "type": "string", "minLength": 1, "maxLength": 120 },
    "audioFile": { "type": "string", "location": "staged directly inside the run folder" },
    "backgroundId": { "type": "string", "source": "assets.json#backgrounds", "description": "Selects one exact checksum-registered fixed background. Sisters Room is the main/default choice." },
    "sequencePreset": {
      "const": "talk-to-camera",
      "mutuallyExclusiveWith": "sequence",
      "description": "Default dialogue preset. Init measures the supplied audio, holds the registered neutral-listening body for that exact duration, and requires Cherry lip-sync."
    },
    "durationFrames": {
      "type": "integer",
      "minimum": 1,
      "maximum": 1800,
      "createdBy": "runner init only when sequencePreset is talk-to-camera",
      "userSupplied": false
    },
    "planningTranscriptSha256": {
      "type": "string",
      "format": "sha256",
      "requiredWhen": "audio-backed gesture sequence or shaz-body-language-performance-v1",
      "description": "SHA-256 printed by the preflight npm run transcribe command. Init regenerates the transcript and rejects stale choreography. Talk to Camera does not require a planning hash."
    },
    "transcript": {
      "type": "object",
      "requires": ["audioFile"],
      "createdBy": "runner init from the bundled whisper.cpp source and English base Q5_1 model",
      "properties": {
        "file": { "const": "transcript.json" },
        "sha256": { "type": "string", "format": "sha256" },
        "receiptFile": { "const": "transcription-receipt.json" },
        "receiptSha256": { "type": "string", "format": "sha256" },
        "sourceAudioSha256": { "type": "string", "format": "sha256" },
        "language": { "const": "en" },
        "segmentCount": { "type": "integer", "minimum": 0 },
        "wordCount": { "type": "integer", "minimum": 0 }
      }
    },
    "lipSync": {
      "type": "object",
      "requires": ["audioFile", "backgroundId"],
      "createdBy": "runner init from bundled Cherry WASI output or a supplied Cherry TSV",
      "properties": {
        "engine": { "const": "cherry-lip-sync" },
        "engineVersion": { "const": "0.1.0" },
        "execution": { "enum": ["node-wasi-preview1", "external"] },
        "cueSource": { "enum": ["bundled-wasi-engine", "supplied-tsv"] },
        "cueFile": { "type": "string", "format": "validated Cherry A-K/X timestamp TSV staged directly inside the run folder" },
        "cueSha256": { "type": "string", "format": "sha256" },
        "sourceAudioSha256": { "type": "string", "format": "sha256" },
        "fps": { "const": 24 },
        "filterSingleFrames": { "type": ["boolean", "null"], "description": "true for bundled WASI generation; null when supplied-cue filter provenance is unknown" },
        "engineManifestSha256": { "type": "string", "format": "sha256", "requiredWhen": "cueSource is bundled-wasi-engine" },
        "engineModuleSha256": { "type": "string", "format": "sha256", "requiredWhen": "cueSource is bundled-wasi-engine" }
      }
    },
    "sequence": {
      "type": "array",
      "minItems": 1,
      "maxItems": 24,
      "items": {
        "required": ["poseId"],
        "properties": {
          "poseId": { "type": "string", "source": "poses/index.json" },
          "holdFrames": { "type": "integer", "minimum": 0, "maximum": 120, "default": 12 },
          "gapFrames": { "type": "integer", "minimum": 0, "maximum": 24, "default": 0 },
          "anchor": {
            "type": "object",
            "requiredWhen": "audio-backed entry whose poseId is not neutral-listening",
            "properties": {
              "wordId": { "type": "string", "description": "Exact word ID from the planning transcript" },
              "label": { "type": "string", "minLength": 1, "maxLength": 120 },
              "frame": { "type": "integer", "minimum": 0, "description": "Entry start frame and round(word.startMs × 24 / 1000)" }
            }
          }
        }
      }
    }
  },
  "initialization": {
    "semanticPreflight": "Before writing a gesture sequence, run npm run transcribe -- --audio=/absolute/path/audio --output=/absolute/path/transcript.json, read its text plus word timestamps, and copy its printed transcriptSha256 into planningTranscriptSha256. This preflight makes no provider call.",
    "transcriptDefault": "Every audio-backed runner init regenerates transcript.json and transcription-receipt.json locally, binds both to the staged audio, and adds transcript to the staged input. Source inputs must not supply their own transcript object.",
    "audioDefault": "When --audio is supplied for a shaz-sequence-input-v1 input with backgroundId, init generates Cherry cues locally with the bundled 0.1.0 WASI engine and adds lipSync to the staged input.",
    "suppliedCueOverride": "--lipsync-cues=/absolute/path/cherry.tsv stages and validates that exact cue file instead of generating one.",
    "explicitOptOut": "--lipsync=off stages audio without creating lipSync.",
    "talkToCameraPreset": "With sequencePreset talk-to-camera, init derives durationFrames from the measured audio and resolves one gap-free neutral-listening body hold. The user supplies no sequence or frame math, and lip-sync cannot be disabled.",
    "unsupportedMode": "Automatic Cherry cue generation and application do not apply to shaz-body-language-performance-v1. Local transcription applies to every audio-backed init, including performance inputs.",
    "performanceBackgroundDefault": "shaz-body-language-performance-v1 accepts an optional backgroundId from the same registry and resolves an omitted value to assets.defaultBackgroundId."
  },
  "maximumOutputFrames": 1800,
  "notes": [
    "Pose IDs are registry entries, never arbitrary filesystem paths.",
    "Every action duration comes from its recipe; holds and any intentional white separators are explicit input fields.",
    "Actions are contiguous by default with no inserted separator frames. Set a positive gapFrames value only when a white separator is intentionally desired; zero does not imply a polished transition.",
    "No finished artist-rendered frame may be supplied as an input.",
    "Talk to Camera is a composition preset, not a fifteenth pose. It reuses the checksum-bound neutral-listening body and changes only the Mouth drawing through the validated Cherry track.",
    "An audio-backed sequence must contain exactly round(audio duration × 24) frames and use a checksum-registered background.",
    "Lip-sync changes only the Mouth drawing. Init generates cues locally by default for an audio-backed sequence, accepts a supplied exact-audio TSV, or disables mouth animation only when --lipsync=off is explicit.",
    "The transcript is planning evidence, not animation data. It may inform pose choice and timing, but it cannot change rig controls, drawings, camera, audio, or rendering by itself.",
    "Every expressive entry in an audio-backed sequence is tied to one exact transcript word and frame. Init rejects the complete plan when the regenerated transcript SHA does not match planningTranscriptSha256.",
    "The Photo Zone background is a fixed clean background in this release. Its reserved supporting-media bounds are metadata only; no overlay input is accepted or rendered."
  ]
}
Quality gatesquality.json
{
  "schemaVersion": 1,
  "automaticGates": [
    "source Xstage provenance",
    "compiled asset checksums",
    "pose recipe registry checksum",
    "native-frame shoulder-to-sleeve-to-hand topology, or exact registered part-replacement tuple, component, exclusivity, and placement contract on replacement frames",
    "native-frame hand-to-sleeve cuff ownership and contact",
    "native-frame hand-to-sleeve and hand-to-head proportion limits",
    "per-frame clipping",
    "per-frame finished layer order, including recipe-declared native arm crossovers, with hidden construction-arm and stray back-bang rejection",
    "per-frame facial pop limits",
    "pose-specific consecutive-identical-frame limit",
    "exact prop presence",
    "1280x720 H.264 yuv420p at 24 fps",
    "exact timeline duration",
    "AAC presence and audible volume for audio-backed runs",
    "registered background id, exact background checksum, and fixed camera for audio-backed runs",
    "waist-up hoodie continuation below the bottom frame edge and clear horizontal margins across every used action",
    "local transcript checksum, source-audio checksum, canonical PCM checksum, ordered word timings, whisper.cpp source and build-plan checksums, locally compiled binary checksum, model checksum, and zero-provider receipt",
    "gesture planning transcript SHA plus an exact transcript word and 24 fps frame anchor for every expressive audio-backed sequence entry",
    "lip-sync cue source, audio checksum, cue checksum, bundled engine artifact checksum when locally generated, engine version, five-mouth mapping, and final resting mouth when lip-sync is requested",
    "talk-to-camera preset provenance, exact measured-audio frame count, exclusive neutral-listening body ownership, zero gaps, and contact-sheet samples at real mouth changes"
  ],
  "humanReview": {
    "required": true,
    "questions": [
      "Do action phases, accents, and holds read as intentional animation without choppy pops or accidental pauses?",
      "Does Shaz remain on-model throughout every action?",
      "Are props and substitution redraws integrated naturally?",
      "Do action boundaries play continuously without unintended white separator flashes?",
      "Would the output be usable in the Shaz animated channel?",
      "When lip-sync is active, do mouth attacks and held shapes follow the actual speech rate without changing body choreography?",
      "Does the mouth remain still during real pauses and finish on the resting drawing?",
      "Does the waist-up crop feel intentional, with no visible lower hoodie boundary that reveals the missing legs?"
    ]
  },
  "attemptLimit": 3
}
Proof reportPROOF-REPORT.md
# Animate Shaz: proof report

Proof baseline: Format 0.4.0

The package can transcribe local English audio with word timing, turn that audio into a complete Shaz talking scene, animate the mouth with five real rig drawings, place Shaz over four built-in rooms, and deliver an inspected 1280×720 video. The 12-second showcase is a strong first draft, not a claim of final creative approval.

## What is safe to use

Use `neutral-listening`, `present`, `think`, `aha`, `point`, and `confident` as the current default building blocks.

Eight other recipes remain in the registry because they are useful for engineering and repair: `shrug`, `key-point`, `excited-celebration`, `point-at-screen`, `look-at-phone`, `facepalm-frustrated`, `arms-crossed-skeptical`, and `phone-use-sequence`. They all need a fresh complete visual review before use in a user video.

Registry checks prove that a recipe can run and pass mechanical rules. They do not prove that the pose looks good.

## Runtime facts

- Source Xstage SHA-256: `507e8b0fa7b95d36b9429671b6b6a9ffa3dd77f5c559b84eb2b49add04512fca`
- Compiled rig assets verified: 210
- Registered pose recipes: 14
- Automated tests: 125 passing, including reproducible package bytes, bundled-engine parity, transcript and cue provenance, tamper rejection, transcript-anchored choreography, local-only audio ingress, cache-link rejection, audio-backed rendering, fixed-stage framing, the duration-derived Talk to Camera preset, the exact four-background registry, and a smoke-fixture guard that excludes needs-review poses
- Registry inspection: all 14 registered actions, 461 recipe frames, zero mechanical failures
- Official smoke: Present and Confident only, 40 frames over 1.666667 seconds; validation, rendering, inspection, and finalization pass without using a needs-review action
- Provider calls: 0
- Cost: $0
- Finished artist-rendered frames used by the runtime or as generation input: false

## Talking-scene proof

- Historical Format 0.2.0 output: 288 frames, 12.0 seconds, 1280×720, H.264 + AAC
- Exact video SHA-256: `59cef6b0910a9d7f8dfe342c0602e8f1921ec6c837fe0fb26c8d5510fd1d2edf`
- Lip-sync: 100 Cherry WASI cues mapped to five recovered mouth drawings
- Source audio SHA-256: `37648e2e0b4c37d22ec529b4301c8abd9435f92ce641c804e521e5b3bdd23f1b`
- Cue SHA-256: `6a6d5604461dfe9aadc40ab3bf5f7b5171c64e674c95384c58275294ee32820d`
- All six used pose recipes passed inspection
- The full audio and video streams decoded successfully
- The run used Sisters Room, no camera movement, and no provider calls

The user chose this exact 12-second result for the main Repo video after calling it very good for a first draft. Its saved `human-review.json` is still `pending`, so it remains a working proof rather than a final creative certification.

## Local transcription status

Development calibration is green: the bundled whisper.cpp source and English model compiled a roughly 2.3 MB arm64 helper locally, without downloading or running a prebuilt native executable. A 30-second Shaz dialogue clip produced 104 timed words and 12 segments in about 1.5 seconds on this Mac. Repeating the transcription produced byte-identical canonical JSON, SHA-256 `35a55600eafd90cb352c01e43e477bcddd945eca6dc21cf757049b059188af7b`.

The sealed-package proof is also green. A fresh blind operator received only the Format ZIP and two different 30-second WAV files. It read each transcript, chose three sparse gestures from the actual words, bound those gestures to six exact word IDs and frames, generated Cherry mouth cues, and completed `init → validate → render → inspect` twice. Both 720-frame H.264 + AAC videos passed with zero failures and decoded completely. One plan used Think, Point, and Confident for a reflective puppy story; the other used Present, Point, and Ah-ha for a family story. The different choices came from different transcript meanings, not volume peaks.

The quarantined extraction compiled its own arm64 helper without a Gatekeeper prompt or quarantine removal. No Whisper executable or model was downloaded, and every transcription/render step reported zero provider calls and $0 cost. A one-time lockfile-pinned `npm ci` still needed network access for Sharp; operation after installation stayed local. Exact transcripts, plans, word anchors, engine receipts, media hashes, and the no-approval boundary are recorded in `evidence/local-transcription-proof.md`. Human review of the two videos remains pending.

## Body-language proof history

- The five recreated artist-authored actions produced 242 frames over 10.083333 seconds. Exact video SHA-256: `8a5183154aaefb9a844d3bd8be48170e7576986f0d0717de5243057c2ea435ae`.
- Present, Think, Ah-ha, Point, and Confident passed all 170 recipe frames with zero mechanical failures.
- Present, Think, Ah-ha, and Point have direct checksum-bound user approval.
- Confident passed synchronized frame comparison plus a fresh-package render and inspection. The user explicitly delegated its final visual decision while away; the agent watched both the exact approval artifact and the complete five-action video through `ended=true`. This does not claim that the user personally watched Confident.
- The legacy Format 0.1.0 ten-action golden has 504 frames over 21.0 seconds at 1280×720, H.264, yuv420p, and 24 fps. It is retained as pending-review history, not current certification.
- The current ten-action fixture has 512 frames over 21.333333 seconds under the pre-Cherry 0.1.2 runtime. The legacy golden does not represent it.
- The current alternate fixture has 191 frames over 7.958333 seconds, with four actions in a different order and different holds.
- A historical blind ZIP-only proof produced 334 frames over 13.916667 seconds from seven independently chosen actions. It proves that an older archive operated on its own; it does not certify the current ZIP.
- Historical structural run `anatomy-v8-release` produced 173 frames over 7.208008 seconds. Exact video SHA-256: `bcf3556ffde53beb7e9efe989bd7e26655b0a2f3a23a5e80ed63f334d0edc9f9`.

`anatomy-v8-release` passed its mechanical checks and received delegated approval at the time. The user later saw and rejected its visible poses. That direct rejection supersedes the earlier delegated acceptance. Keep the run as engineering history; do not use it as a showcase or creative release gate.

## Talk to Camera proof

`sequencePreset: "talk-to-camera"` is a composition preset, not another pose. Initialization measures the supplied audio, records `durationFrames`, resolves one gap-free `neutral-listening` hold for the exact length, and requires Cherry cues tied to that audio. The user supplies no `sequence`, `holdFrames`, or frame calculation.

Validation, render, inspection, and delivery receipts carry the preset ID. The contact sheet samples real mouth-change frames. The preset reuses the existing one-frame neutral anchor and adds no body controls, props, camera movement, or renderer branch. Explicit gesture sequences still work and keep their 120-frame-per-entry hold limit.

A sealed Format 0.2.1 extraction passed all 110 packaged tests. The minimal preset then ran against two meaningfully different files: a 2.843-second OGG and a 12-second WAV. Both timelines came from the audio with no user frame math. Both used only `neutral-listening`, created local Cherry cues, produced H.264 + AAC, passed inspection with zero failures, and decoded completely. Exact receipts live in `evidence/talk-to-camera-preset-proof.md`. Creative approval remains a separate decision.

## Background proof

`assets.json` contains exactly four opaque 3840×2160 RGB PNG backgrounds and names `sisters-room` as the default. An audio-backed sequence or Talk to Camera run names a registered ID. The separate semantic performance mode accepts the same IDs and resolves an omitted choice through the manifest. Every background stays behind the same fixed stage and character renderer.

- Living Room is an exact flattened composite of the supplied PSD.
- Photo Zone hides only the visible source map-art layer and preserves the clean purple striped wall beneath it. Its recorded supporting-media bounds are reserved for a future feature; the current Format accepts and renders no overlay there.
- Pure White is an exact generated `#FFFFFF` canvas.

A sealed extraction passed all 111 packaged tests and `npm run check`, contained exactly the four registered background files, and rendered the same Talk to Camera audio, timing, and cues twice: once in Living Room and once in Photo Zone. Both outputs were 68-frame H.264 + AAC videos. Both decoded completely and passed inspection with zero failures. Their receipts record different background IDs and checksums while audio, cues, pose, holds, camera, and mouth histogram remain identical. Exact receipts live in `evidence/background-library-proof.md`. Creative approval remains separate.

## Cherry proof

When an audio-backed `shaz-sequence-input-v1` run starts, the bundled Cherry Lip Sync 0.1.0 WASI module creates cues locally by default. A Cherry TSV from the exact audio can be supplied instead, and `--lipsync=off` is the explicit audio-without-mouth-motion path.

The package contains no native Cherry executable and makes no provider call. The cue-only command and sequence initializer use the same bundled engine. `renderRigFrame` remains the only character renderer.

The separate `shaz-body-language-performance-v1` mode remains body-language-only. Its audio controls duration and gesture scheduling, not mouth drawings.

For the historical Format 0.2.0 package proof, a sealed extraction installed, passed all tests, passed `check`, passed registry inspection, and passed smoke without a supplied cue TSV. Two different speech inputs then completed `init → validate → render → inspect` through the bundled engine. Both videos decoded completely and reported zero mechanical failures. Exact receipts live in `evidence/bundled-cherry-wasi-proof.md`. This proves the bundled cue engine; it is separate from the later Talk to Camera preset proof.

## Publication boundary

The rich Repo page is testable, but the Format stays off the main Discovery shelf until the exact talking-proof checksum receives final creative approval.

The downloadable ZIP intentionally contains the pinned whisper.cpp source archive and English model. It does not contain the original Toon Boom archive, source PSDs, finished artist renders, agent runs, downloads, `node_modules`, or a downloaded native Whisper executable.
ProvenancePROVENANCE.md
# Where the pieces came from

## Shaz's rig

The runtime was recovered from the Toon Boom Xstage project supplied by the user.

Source Xstage SHA-256:

`507e8b0fa7b95d36b9429671b6b6a9ffa3dd77f5c559b84eb2b49add04512fca`

The original archive is not in the package. `rig-v2/runtime.json` records the recovered hierarchy, channels, drawing exposures, camera conversion, and compositing plan. `rig-v2/assets/receipt.json` records each compiled drawing asset and its source and output checksums.

Finished artist-rendered video frames were used only as human comparison material during development. They are not packaged, are not runtime assets, and were not used as generation input for the six new actions. Every recipe and render receipt says `artistRenderedFramesUsed: false`, and validation rejects a recipe that does not.

The phone is a small purpose-built non-limb prop. Most substitutions use existing compiled rig drawings with fixed checksums.

Crossed Arms is the one drawing exception. After native-rig anticipation, it uses a fixed arm-only destination drawing derived from the crossed-arms pose artwork supplied by the user. The file and placement are locked. The runtime still draws the original head, hair, face, torso, collar, strings, and pocket. Both native arm chains disappear on the same frame that the replacement appears, so the two versions cannot paint on top of each other. No finished animation frame or image-generation output is packaged as character artwork.

That origin record does not make Crossed Arms creatively approved. Its current recipe still needs fresh visual review.

The character assets are packaged for the owner's authorized Wiggly workflow. This package does not give third parties permission to redistribute or commercially use the character or source art.

## Backgrounds

The package includes four opaque 3840×2160 RGB PNG backgrounds. `assets.json` records each output checksum, label, dimensions, use, source operation, and source checksum when a PSD exists.

- **Sisters Room** is the default. It comes from `BG (8) Sisters room.psd`, SHA-256 `5ad1d74940954256925905428fb945bd07ecf4a22d104ff42c55696caa6c5566`.
- **Living Room** comes from the embedded composite in `BG (4) living room.psd`, SHA-256 `1da1c55eeccea94b49b14634b8167d2aeeee9be5618fff94d166d95595e3bd3d`. Boundary alpha was flattened onto white before deterministic RGB encoding.
- **Photo Zone** comes from `BG (22) map.psd`, SHA-256 `2666ddcf35837d74dc3e80803e138331a0f61f0ea46f78a42c535037a58eeb19`. Only the visible `Layer 4` map artwork was hidden. The PSD's already-hidden smart-object layer named `map` stays hidden. The cleared wall is a fixed background; its recorded bounds do not turn on a supporting-media renderer.
- **Pure White** is a deterministic generated `#FFFFFF` RGB canvas with no source PSD.

The source PSD files are not packaged. Every room uses the same character renderer, fixed camera, and character transform.

## Cherry Lip Sync

Beginning with Format 0.2.0, the package includes Cherry Lip Sync 0.1.0 as a `wasm32-wasip1` module. Cherry creates A-K/X speech cues before Shaz is rendered. It does not render Shaz and contains no Shaz artwork.

- Upstream repository: <https://github.com/amberwhitehead/cherry-lip-sync>
- Immutable tag: `v0.1.0`
- Immutable commit: `ab3e68a8e2d38fc72d1672c450478dff7710bc14`
- Source archive SHA-256: `cf16c5bf5fdeed18a96240e74144287f80957c1c1461891eb74585ac3ab94bfc`
- Bundled WASI module SHA-256: `1bf5730acc7a81b1f0b6c818a9068001b9e9a797a7eb990ae091eb1b01603382`
- License: `MIT OR Apache-2.0`

The source, build receipt, dependency-only patch, parity fixture, module checksum, `LICENSE-MIT`, `LICENSE-APACHE`, `NOTICE`, and upstream README are under `vendor/cherry-lip-sync/v0.1.0/`.

The build patch marks two unused build-time dependencies optional so the command-line target can compile for WASI. It changes no Rust source, model, inference, audio decoding, command-line behavior, or cue-generation logic. The fixture cues produced by the bundled module are byte-for-byte identical to the upstream macOS arm64 0.1.0 output.

The package contains no native Cherry executable. Node runs the module with only a private temporary scratch folder open to WASI, so the downloaded kit never asks the user to bypass macOS Gatekeeper. Rust is needed only to rebuild the module from source, not to use the package.

## Local transcription

Beginning with Format 0.4.0, the package includes whisper.cpp 1.9.2 source and the English `base.en` Q5_1 model. It compiles a small arm64 helper locally with Apple Clang and Accelerate, then uses that helper to create text and word timestamps before body-language planning. The helper is build output, not a downloaded executable, and it is excluded from the distributable ZIP.

- Upstream repository: <https://github.com/ggml-org/whisper.cpp>
- Immutable tag: `v1.9.2`
- Immutable commit: `306c88f4d1286aec1bf96e544632897886af5501`
- Source archive SHA-256: `a6abd064fcca8b85e794d205abf328c522e9451db43a3eadc178b883b7d0e9cd`
- Bundled model: `ggml-base.en-q5_1.bin`
- Model repository commit: `5359861c739e955e79d9a303bcbc70fb988958b1`
- Model SHA-256: `4baf70dd0d7c4247ba2b81fafd9c01005ac77c2f9ef064e00dcf195d0e2fdd2f`
- Engine and model license: MIT

`vendor/whisper.cpp/v1.9.2/VENDOR-MANIFEST.json` binds the source, model, build plan, and license files. `BUILD-PLAN.json` fixes the compiler inputs, architecture, minimum macOS version, and Accelerate linkage. No user audio leaves the machine, and no Deepgram or other hosted transcription service is used.
Advanced execution detailsCommands and the published proof receipt, available when you need them.
Read from package.json scripts
  1. $ npm run check
  2. $ npm run inspect:registry
  3. $ npm run smoke
  4. $ npm run transcribe -- --audio=/absolute/path/audio.wav --output=/absolute/path/transcript.json
  5. $ npm run lipsync -- --audio=/absolute/path/audio.wav --output=/absolute/path/cherry.tsv
  6. $ npm run init -- --run=episode-01 --input=/absolute/path/input.json --audio=/absolute/path/dialogue.wav
  7. $ npm run validate -- --run=episode-01
  8. $ npm run render -- --run=episode-01
  9. $ npm run inspect -- --run=episode-01
  10. $ npm run finalize -- --run=episode-01

Published proof receipt

Kit version
0.4.0
Demo made with
0.2.0
Video SHA
59cef6b0910a9d7f
Frames
288
Output
1280 × 720 · 24 fps
Audio
AAC · -15.6 dB
Lip-sync
100 cues · 5 mouths
Artist frames
Excluded

This 0.2.0 proof rendered from a fresh download and passed validation plus the audio, video, and rig checks. The exact video checksum is recorded above. Human creative review is still pending, so this is a strong first draft—not a finished performance. User feedback: praised-as-good-first-draft; exact-checksum final creative approval not claimed.

Run it with a coding agent

Know the run before you start.

Add a voice track and pick a room. The kit reads what Shaz is saying on your Mac, lip-syncs the mouth, and helps place an artist-reviewed gesture on the words that deserve one.

Typical run

Plan + checks$0 · under 1 min
Render on your Mac$0 · about 1-5 min
Watch + review$0 · about 1-3 min

$0 in service fees, usually 2-8 min

You provide

One dialogue audio file · One built-in background; Sisters Room is the default · Talk to Camera for ordinary speech, or a short sequence made from the five artist-reviewed gestures · Hold and pause timing only when using a custom gesture sequence

Output

One 1280 × 720 talking-scene MP4 in the room you chose, plus the checks and review record used to deliver it