Bits, bytes and the range of numbers
A bit is one switch: 0 or 1. A byte is 8 bits. Each place has a weight, doubling as you go left: 1, 2, 4, 8, 16, 32, 64, 128.
With n bits there are 2n different patterns. If all of them are used for non-negative numbers (unsigned), the range is 0 to 2n β 1. For 8 bits that is 0 to 255. For 16 bits it is 0 to 65 535.
Example: 00101101 = 32 + 8 + 4 + 1 = 45.
Two's complement: storing negative numbers
We need some patterns to mean negative numbers. In two's complement the top bit (the most significant bit) has a negative weight. In 8 bits it is β128 instead of +128.
- If the top bit is 0, the number is 0 to 127.
- If the top bit is 1, the number is β128 to β1.
So the signed 8-bit range is β128 to +127. In n bits it is β2nβ1 to 2nβ1 β 1. There is one more negative number than positive ones, and only one zero.
How to write βx
- Write x in binary.
- Flip every bit (0 to 1, 1 to 0).
- Add 1.
Example: β5. 5 = 00000101. Flip: 11111010. Add 1: 11111011. Check: β128 + 64 + 32 + 16 + 8 + 2 + 1 = β5.
Why computers like it
The same adder circuit adds positive and negative numbers. 5 + (β5) = 00000101 + 11111011 = 1 00000000. The extra 9th bit is dropped, leaving 0. Subtraction becomes adding the negative.
Overflow: when the answer does not fit
A byte has only 256 patterns. If a result is outside the range, the extra bit is lost and the number wraps round like an odometer. This is overflow.
- Signed 8-bit: 127 + 1 = β128.
- Unsigned 8-bit: 255 + 1 = 0.
- Unsigned 8-bit: 200 + 100 = 300 β 256 = 44.
Signed overflow can be spotted: two positive numbers add to a negative one, or two negative numbers add to a positive one. Real programs avoid it by using more bits (16, 32, 64) or by checking the result first.
Bitwise operations and shifts
Bitwise operations treat a number as a row of bits and work on each column separately.
| Operation | Rule | Example (8 bits) |
|---|---|---|
| AND (&) | 1 only if both bits are 1 | 1101 & 1011 = 1001 |
| OR (|) | 1 if at least one bit is 1 | 1101 | 1011 = 1111 |
| XOR (^) | 1 if the bits are different | 1101 ^ 1011 = 0110 |
| NOT (~) | flip every bit | ~00000101 = 11111010 |
Shifts
- Left shift << k: move bits k places left, fill with 0. It multiplies by 2k (if nothing falls off the top). 5 << 3 = 40.
- Right shift >> k: move bits right. It divides by 2k and drops the remainder. For negative signed numbers the top bit is copied in (arithmetic shift), so β5 >> 1 = β3.
Uses: AND with a mask picks out some bits, OR sets bits, XOR flips chosen bits, and shifts multiply or divide by powers of 2 very fast.
Floating point: storing decimals
To store numbers like 3.14 or 0.000001 the computer uses floating point, like scientific notation in binary. The bits are split in three parts:
- Sign (1 bit): 0 is positive, 1 is negative.
- Exponent: says how far the point moves (a power of 2). It is stored with a bias, so it can be negative too.
- Mantissa (fraction): the digits of the number, written as 1.xxxx (the leading 1 is not stored).
Value = (β1)sign Γ 1.mantissa Γ 2exponent β bias.
Our 3D uses a tiny 8-bit version: 1 sign, 3 exponent bits (bias 3), 4 mantissa bits. 0 100 1000: exponent 4 β 3 = 1, mantissa 1.5, value 1.5 Γ 2 = 3. Real computers use 32 bits (float: 1 + 8 + 23) or 64 bits (double: 1 + 11 + 52).
Why 0.1 is not exact
0.1 in binary never ends: 0.000110011β¦ Only a fixed number of bits can be kept, so the stored value is a very close guess. That is why 0.1 + 0.2 shows 0.30000000000000004. Never test floating point numbers with ==; check if they are close enough. Very large exponents give infinity and invalid results give NaN (not a number).
Key formulas and definitions
- Unsigned n bits: 0 to 2^n β 1
- Signed n bits (two's complement): β2^(nβ1) to 2^(nβ1) β 1
- Negative of x: flip all bits of x, then add 1
- x << k = x Γ 2^k; x >> k = floor(x Γ· 2^k)
- Float value = (β1)^sign Γ 1.mantissa Γ 2^(exponent β bias)
- float32 = 1 + 8 + 23 bits; double = 1 + 11 + 52 bits
Worked examples
1. Find the value of 00101101 as an unsigned 8-bit number.
Weights of the ON bits: 32 + 8 + 4 + 1 = 45.
2. Write β20 as an 8-bit two's complement number.
20 = 00010100. Flip: 11101011. Add 1: 11101100. Check: β128 + 64 + 32 + 8 + 4 = β20.
3. What is 11110110 as a signed 8-bit number?
Top bit is 1, so it is negative: β128 + 64 + 32 + 16 + 4 + 2 = β10.
4. A signed 8-bit variable holds 120. We add 10. What is stored?
120 + 10 = 130, which is above 127. It wraps: 130 β 256 = β126. In bits: 10000010 = β128 + 2 = β126.
5. Find 13 XOR 6 and 5 << 3.
13 = 1101, 6 = 0110. XOR gives 1011 = 11. And 5 << 3 = 5 Γ 8 = 40 (00101000).
6. In the 8-bit mini-float (1 sign, 3 exponent bits with bias 3, 4 mantissa bits) find the value of 0 011 1100.
Exponent = 3, so 3 β 3 = 0. Mantissa 1100 = 12/16 = 0.75, so 1.75. Value = 1.75 Γ 2^0 = 1.75.
7. What is β5 >> 1 on a signed 8-bit number?
β5 = 11111011. Arithmetic right shift copies the top bit: 11111101 = β3 (the result rounds down, toward βinfinity).
Common mistakes
- Reading the top bit as +128 when the number is signed. In two's complement it is β128.
- Forgetting to add 1 after flipping the bits. Flip alone gives βx β 1.
- Thinking overflow gives an error message. Most languages and hardware silently wrap round.
- Comparing decimals with == in code. 0.1 + 0.2 == 0.3 is false; compare with a small tolerance.