📘 CodingMarble Learn

Sound and Voice Content: Expression, Editing and Production

Sound and voice can say things text cannot: feeling, rhythm, place and mood. On a computer, sound is stored as a long list of numbers (a waveform). Editing changes those numbers: cut, fade, change volume, remove noise, mix tracks. Producing content means planning, recording, editing, mixing, exporting and sharing.

🎬 Step-by-step story

  1. Speak. The air shakes, and the shaking is drawn as a wave. The wave is how a computer sees sound.
  2. Speak louder. The wave becomes taller. Taller wave means louder sound (more volume).
  3. Make the voice thin and high. There are more waves in the same time. More waves means higher pitch.
  4. Edit: the red part is extra. We cut it away. The rest of the sound is not harmed.
  5. Produce: voice, music and a sound effect sit in separate tracks. We mix them into one finished sound.
  6. Free play: change volume, pitch and the cut. Watch the wave and the caption below it.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

What is a waveform really?

It is a picture of the measured numbers of the sound: time goes left to right and the height shows how strong the sound is at that moment.

Is a louder sound also a higher sound?

No. Loud is a taller wave. High is more waves in a second. A shout can be low and a whisper can be high.

Why does a thin voice make more waves?

A thin, high voice shakes the air faster, so more waves fit in each second. A deep voice shakes slowly.

If I cut a part, does the rest change?

No. Cutting removes only the selected part. The other waves stay the same, so keep the raw copy in case you cut too much.

Why do we need separate tracks?

On separate tracks you can change voice, music and effects one by one, then mix them. If they were joined you could not fix one without hurting the others.

1. Expression with sound and voice

Sound can tell us things that words on paper cannot. The same sentence "I am fine" can sound happy, tired or angry. That difference comes from the voice: how loud, how fast, how high, and where the speaker pauses.

Four kinds of sound content

Why choose sound?

Sound works when eyes are busy (driving, cooking), for people who cannot see well, and for feeling. It is weak for maps, tables and long numbers, where a picture is better. Good voice work is clear: not too fast, one idea at a time, and no background noise.

2. How sound becomes numbers

Real sound is a smooth wave in the air. A microphone turns it into a changing electric signal. The computer then samples it: it measures the height of the wave many times each second and writes each measurement as a number.

More samples and more bits sound better but make a bigger file. One minute of CD-quality stereo sound is about 10 MB.

3. Editing sound and voice

An audio editor shows the waveform. You can select a part and change it.

Non-destructive editing keeps the original file safe, so you can undo. Always keep a copy of the raw recording.

File types

WAV keeps every number (big, best quality). MP3 and AAC throw away sounds people hardly hear to make small files (compressed, "lossy"). Record and edit in WAV, then export to MP3 to share.

4. Producing sound and voice content

Producing means making a finished piece for an audience, for example a podcast, an audio guide, a radio ad or the sound for a video. Follow these stages.

  1. Plan: who will listen, what do they need, how long? Write a script and a list of sounds.
  2. Record: use a quiet room, keep the microphone a hand-width from your mouth, do a test, and record extra takes.
  3. Edit: cut mistakes, clean noise.
  4. Mix: balance voice, music and effects.
  5. Export: save as MP3 or AAC with a clear file name.
  6. Check and share: listen on a phone and on a speaker, then publish.

Rules to follow

Ask permission before using another person's voice. Use music that is free to use or that you have a licence for. Credit the people who helped.

Try it

At home: record yourself reading the same two lines three times: calm, excited, sad. Listen. Which parts of your voice changed? Now use a free audio editor to cut the silence at the start and add a 1-second fade-out. Compare before and after.

In the 3D: before you move the Volume slider, guess what will happen to the wave. Then check.

Key formulas and definitions

Worked examples

1. In a waveform, the first clip has tall waves and the second clip has short waves. Which is louder?

The first clip. Height of the wave is volume, so tall waves are louder.

2. Clip A has 200 waves in one second and clip B has 400 waves in one second. Which has the higher pitch?

Clip B. More waves per second means higher pitch. B is twice as high as A.

3. You record 10 seconds of mono sound at 8,000 samples per second and 16 bits. How big is the raw file?

Bits = 8,000 × 16 × 1 × 10 = 1,280,000 bits. Bytes = 1,280,000 ÷ 8 = 160,000 bytes, about 160 KB.

Common mistakes

Practice quiz

1. Which part of a waveform shows volume?
2. Which file type keeps all the sound numbers without loss?
3. A fade-out does what?
4. What is the first stage of producing a podcast?
5. Sampling rate means:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is sound and voice content?

Anything made of sound for an audience: speech, music, sound effects, podcasts, audio guides and the sound track of a video.

What is the difference between volume and pitch?

Volume is how loud (height of the wave). Pitch is how high or low (how many waves per second).

Which is better for sharing, WAV or MP3?

MP3 is small, so it is better for sharing. WAV is better for editing because it keeps full quality.

Where this is taught

Japan高校(専門学科)1〜3年Content Production and Publishing

Learn first

Learn next

Related lessons

All Computer Science lessons