4-Layer Editorial System
Arena's core differentiator is a 4-layer editorial system that produces standalone, high-quality clips while other AI tools output incomplete thoughts. Every clip passes through all four layers.
Layer 1: Hook & Moment Detection
The engine scans the full transcript in sliding 2-minute windows, detecting approximately 40 "seeds" — compelling moments like claims, insights, hooks, stories, and conversational peaks. Seeds are deduplicated by time proximity and text similarity to produce ~25 raw candidate windows.
# Enable layer export to see Layer 1 output
arena process video.mp4 --export-layers
# Check: output/editorial_layers/layer1_moments.jsonLayer 2: Thought Boundary Expansion
Raw clip windows are expanded outward to natural thought boundaries. Each candidate is analyzed for its premise (setup/context), claim (the core point), and resolution (conclusion or punchline). The boundaries are adjusted so clips capture the complete thought arc — premise to resolution.
This prevents the common AI clipping problem of clips that start or end mid-sentence or mid-thought.
Layer 3: Standalone Quality Gate
The strictest layer. A validation agent checks whether each clip can stand completely alone on a social media feed without surrounding context. It checks for:
- Unresolved pronouns ("it", "this", "that" without clear referents)
- Missing subject introductions
- Dangling references to earlier conversation
- Incomplete premises or resolutions
- Structural issues requiring significant adjustments
Each clip receives a standalone score from 0–10. Only clips scoring 7+ pass. Typical pass rate is around 7–10% of candidates.
Debugging rejections
# Export layer data
arena process video.mp4 --export-layers
# Check rejection reasons
cat output/editorial_layers/layer3_validated.json | \
jq '.[] | select(.verdict=="REJECT") | {thought_id, rejection_reason, standalone_score}'
# Check pass rate
cat output/editorial_layers/layer3_validated.json | \
jq '[.[] | select(.verdict=="PASS")] | length'Common rejection reasons
| Reason | Meaning |
|---|---|
duration_constraint | Clip too short or long after boundary adjustment |
missing_premise | Doesn't explain what it's about |
dangling_reference | Uses pronouns without defining them |
incomplete_resolution | Cuts off mid-thought |
structural_issue | Would need 15+ seconds of adjustment to fix |
Layer 4: Viral Packaging
Clips that pass Layer 3 are packaged with professional metadata for social media distribution:
- Magnetic, high-CTR titles
- Detailed keyword summaries
- Platform-specific hashtags
- Hook text for captions
Hybrid scoring
Arena combines two signals to rank clips:
- AI content score — GPT-based semantic analysis of content quality
- Audio energy score — RMS amplitude and spectral centroid analysis detecting enthusiastic delivery
Final score = ai_score × 0.7 + energy_score × 0.3. This ensures clips have both great content and dynamic delivery.
Semantic deduplication
Layer 4 also deduplicates clips using text-embedding-3-small cosine similarity. If multiple clips cover the same topic, the best variant per cluster is selected, preventing repetitive output.
Checkpointing
The CheckpointManager saves intermediate results between layers. If a run fails mid-pipeline (e.g., rate limit or network error), Arena resumes from the last completed layer instead of restarting. Job ID is a hash of the first 500 characters of the transcript.
Cost breakdown
| Layer | Model | Typical cost |
|---|---|---|
| Transcription | Whisper API | ~$0.006/min |
| Layer 1–2 | Configurable (--editorial-model) | ~$0.05–0.20 |
| Layer 3–4 | Always gpt-4o-mini | ~$0.03–0.05 |
| Deduplication | text-embedding-3-small | <$0.01 |
Total for a 30-minute video: $0.15–0.25 with gpt-4o-mini, or $0.40–0.60 with gpt-4o.