Sound as a wave
Sound is air moving back and forth. We draw it as a wave. Two things matter most: how many waves per second (frequency, in hertz, Hz, which we hear as pitch) and how tall the wave is (amplitude, which we hear as loudness). The wave shape gives timbre, the "colour" of a sound. 440 Hz is the note A that orchestras tune to.
Synthesis systems: oscillator, filter, envelope
A synthesizer makes sound with electronics or software. Its basic blocks are:
- Oscillator (VCO): makes a repeating wave. Common shapes: sine (pure, one harmonic), square (hollow, odd harmonics), sawtooth (bright and buzzy, all harmonics), triangle (soft, weak harmonics).
- Filter (VCF): removes some harmonics. A low-pass filter lets low sounds pass and cuts highs above the cutoff. This is subtractive synthesis: start rich, then take away.
- Amplifier and envelope (ADSR): controls loudness over time. Attack = time to reach full volume; Decay = fall to the held level; Sustain = the level held while the key is down; Release = fade after the key is up.
- LFO: a slow oscillator that wobbles pitch (vibrato) or loudness (tremolo).
Other methods: additive (add many sine waves, like the harmonic bars), FM (one oscillator bends another) and wavetable (step through stored waves).
Sampling
A computer cannot store a smooth wave. It measures the height of the wave many times a second. Each measure is a sample.
- Sample rate: measures per second. CDs use 44,100 Hz; video often uses 48,000 Hz. To record a sound with highest frequency f, you need a sample rate above 2f (the Nyquist rule). Human hearing reaches about 20,000 Hz, so 44,100 Hz is enough.
- Bit depth: how many levels each sample can take. 16 bits give 216 = 65,536 levels; 24 bits give far more and a quieter noise floor.
- Too low a rate makes a staircase wave and a false, low tone called aliasing.
A sampler plays recorded sounds back at different pitches from a keyboard. Producers also use samples in loops, drum kits and sample libraries. Always respect copyright and ask for permission.
Sound, gesture, text and image
Modern electroacoustic works link sound to other things. Gesture: a sensor, a camera or a touch screen turns hand movement into pitch, volume or filter settings. Text: a poem can be spoken, cut into pieces and turned into sound (speech synthesis, text-to-speech). Image: pictures can be drawn as sound (a spectrogram read backwards) or sound can drive visuals, as in live shows and games. The common thread is mapping: deciding which input changes which sound parameter.
Music online
Digital sound files are shared by streaming and download. Files are made smaller by compression: lossless (FLAC, keeps everything) or lossy (MP3, AAC, drops parts we hear least). Browsers can make sound live with the Web Audio system, so synths can run in a web page. Online platforms also bring copyright and licences (for example Creative Commons) and online collaboration between artists far apart.
Electroacoustic music and its criticism
Electroacoustic music uses electronics and recorded sound as its main material. Early landmarks: musique concrète in Paris in the 1940s (tape recordings of real sounds, cut and changed) and the electronic music studios in Cologne (pure synthesised sounds). Later came computers and live electronics.
To write a critique, listen more than once and answer: What sounds are used (recorded or synthesised)? How does the sound change in time? How is space used (left, right, far, near)? What is the structure? What idea or feeling does it make you think of? Support each point with a moment in the piece, not just "I liked it".
Try it: build and see
In the 3D, (1) pick a saw wave, lower the filter and watch the wave soften; (2) set the sample count to 6, then raise it and see the staircase disappear. At home: clap once near a phone voice recorder, then zoom into the waveform in the app. Predict first, then check: where is the attack, where is the release?
Key formulas and definitions
- Frequency f (Hz) = waves per second; period T = 1 / f
- Nyquist rule: sample rate > 2 × highest frequency
- Number of levels = 2^bits (16 bits → 65,536)
- Data per second = sample rate × bits × channels (bits)
- ADSR = Attack, Decay, Sustain, Release
- Octave up = frequency × 2
Worked examples
1. What is the period of the note A at 440 Hz?
T = 1 / f = 1 / 440 ≈ 0.00227 s, about 2.27 ms.
2. The note A3 is 220 Hz. What is the A one octave above?
An octave doubles frequency: 220 × 2 = 440 Hz.
3. Why is 44,100 Hz enough for music?
Humans hear up to about 20,000 Hz. The Nyquist rule says the sample rate must be more than twice this, 40,000 Hz. 44,100 Hz is above it.
4. How many samples are in 2 seconds of audio at 48,000 Hz (mono)?
48,000 × 2 = 96,000 samples.
5. How many bytes is 1 minute of CD-quality stereo audio?
44,100 samples/s × 16 bits × 2 channels × 60 s = 84,672,000 bits ÷ 8 = 10,584,000 bytes, about 10.6 MB.
6. A pad sound slowly fades in and fades out after the key is lifted. Describe its ADSR.
Long attack (slow rise), medium decay, high sustain, long release (slow fade).
Common mistakes
- Mixing up pitch and loudness: pitch is frequency, loudness is amplitude.
- Thinking a higher sample rate always means higher pitch. It only gives more detail, not a different note.
- Believing that a low-pass filter makes a sound quieter only. It removes high harmonics and makes it duller.
- Confusing sampling (measuring a sound) with synthesis (building a sound).