{
  "contents": [
    {
      "role": "user",
      "parts": [
        {
          "text": "Attachment 1 is a complete English sung advertorial story-song (the inspiration). The other attachments are complete Ukrainian songs telling the same story for a different product (a neck cream). Listen to every attachment end to end. For EACH Ukrainian attachment answer, with timed evidence (attachment, local seconds, quoted words): (1) songOrNarration: does it sound like a SONG (melody, form, phrases that rise and settle, returning material) or like NARRATION set to music (talk-singing, one continuous stream of words)? Give the verdict and the passages that decide it. (2) thoughtsSeparatedByRests: are consecutive thoughts separated by audible rests of the voice (not pauses after every line, but a breath between complete sentences)? Count the clear rests you hear per minute and name three examples. (3) codaIntelligible: in the final product/order block, list every fact you can make out (brand name, order form with name and phone, inspection at the post office, payment after inspection, a 60-day return, a link below) and whether each one has time to land or is crammed. (4) melodyRiseFall: where does the melody clearly go UP (name 3 moments) and where does it come DOWN or settle (name 3)? Is there an audible climax and where? (5) hookReturns: is there a line or refrain that returns; quote it and give each return time. (6) sungThroughout: any passage that slips into spoken recitation (local seconds)? (7) longNotes: which words are held long, and are those the important words? (8) defects[{localSeconds,words,problem,severity}], wordsMisheardOrUnclear[]. Then across all Ukrainian attachments: whichBreathesBest (attachment number, why), whichIsMostSong (attachment number, why), ranking (best first) with one sentence each. Return ONLY JSON (no prose outside JSON), <=1900 words, English. Required keys: mediaAccess(boolean), audioAccess(boolean), attachmentsHeard[{attachment,firstWordsHeard,lastWordsHeard,durationEstimateSeconds}], then the analysis keys listed. No script is supplied: quote only words you actually hear and mark uncertain words with (?). Use attachment number plus LOCAL seconds of that attachment. Do not rate quality with numbers, do not assume any language or attachment is better a priori, do not reward louder mastering, more notes or softer timbre. Identify concrete mechanisms, not taste."
        },
        {
          "text": "Attachment 1: Attachment 1; complete English original song, 194.5s, 56kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 2: Attachment 2; complete Ukrainian song, 197.1s, 48kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 3: Attachment 3; complete Ukrainian song, 233.2s, 48kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        }
      ]
    }
  ],
  "stream": true,
  "generationConfig": {
    "thinkingConfig": {
      "includeThoughts": false,
      "thinkingLevel": "high"
    }
  }
}
