Loading...

HTML Charset

Favicon links tell the browser which tiny image to show on the tab. Charset answers a more fundamental question: how should the browser turn the bytes of your HTML file into letters humans can read? When encoding is wrong, Persian, Arabic, accented Latin, Chinese, and emoji often become mojibake — garbled sequences that look broken even though the rest of the markup is fine.

What charset means

A character set (and its encoding) maps bytes on disk to characters. HTML files are byte streams. Without a correct decoding rule, the browser may guess wrong. Declaring UTF-8 in head is the modern web habit that keeps multilingual text and symbols reliable across editors, servers, and browsers.

Direct answer

Put <meta charset="UTF-8"> near the top of head — ideally within the first 1024 bytes of the file so the browser learns the encoding early. UTF-8 is the encoding you should use for almost all new pages. Also save the HTML file itself as UTF-8 in your editor. Declaration and file save must agree; one without the other still produces broken characters.

Declare, save, and language — different jobs

📊 Encoding versus language versus title

Piece Job
meta charset How bytes decode into characters
Editor save as UTF-8 How your tool writes those bytes to disk
html lang Which human language the page content is in
title text Tab label words (not an encoding switch)

A multilingual UTF-8 demo

One short document shows charset in action. The meta comes first in head, then a title, then body text that mixes scripts.

UTF-8 charset in head

html
HTML Code
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Charset demo</title>
</head>
<body>
<p>Hello — سلام — مرحبا — 你好 — 🙂</p>
</body>
</html>

Reading the sample

  • <meta charset="UTF-8"> sits early in head, before most other content
  • lang="en" names a primary language for tools; it does not replace charset
  • The paragraph includes Latin, Persian, Arabic, Chinese, and an emoji — all fine under UTF-8 when the file is saved correctly
  • If those characters break, check save encoding before rewriting the HTML structure

Why put charset early

Browsers begin reading the file from the top. If charset appears only after a huge block of text or scripts, some characters may already have been interpreted with a wrong guess. Keep the meta near the opening of head. In the starter heads from earlier sessions, charset was intentionally first for this reason.

Mojibake: the symptom of mismatch

Mojibake looks like random punctuation, Latin leftovers, or replacement characters where real words should be. Common causes: declaring UTF-8 while the file was saved as a legacy encoding, or omitting charset and hoping the browser guesses right. Fix the encoding pair — meta plus file save — rather than “fixing” letters by retyping forever.

lang is not charset

The lang attribute on html (and sometimes on specific elements) helps screen readers, hyphenation, and search tools understand language. Charset handles byte decoding. You usually want both: lang for language identity, charset for encoding. Do not remove charset because you already set lang="fa" or lang="ar".

Servers and real deployments (beginner note)

Production servers can also send a Content-Type charset in HTTP headers. For this course, master the HTML meta and UTF-8 save habit first. When a live site disagrees with your meta, ask a mentor to check headers — but most local practice problems are editor save encoding or a missing meta tag.

Practice in your editor

  1. Create a page with <meta charset="UTF-8"> at the top of head.
  2. Add a sentence that includes non-English letters and one emoji.
  3. Confirm the editor status bar says the file is UTF-8, then save.
  4. Open the page in a browser and confirm the characters look correct.
  5. (Optional, carefully) Save a copy as a legacy encoding, reopen, and observe mojibake — then restore UTF-8 so you recognize the failure mode.

Easy mistakes

  • Forgetting meta charset and relying on browser luck
  • Saving the file as a legacy encoding while declaring UTF-8
  • Putting charset meta very late in a huge head
  • Confusing charset with the visible page title or with lang
  • Assuming emoji problems are always “HTML bugs” when encoding is the real cause
  • Fixing mojibake by deleting lang instead of correcting UTF-8 save + meta

Summary

  • Charset declares how to decode the HTML file’s bytes into characters
  • Use UTF-8 for almost all new pages and place meta charset early in head
  • Save the file as UTF-8 so declaration and disk encoding match
  • lang names language; charset names encoding — different jobs
  • Mojibake usually means an encoding mismatch, not a mysterious tag failure

🧠 Test Your Knowledge

Ready to Start

Test Your Knowledge

Challenge yourself with this interactive quiz and see how well you understand the topic

❓
6
Questions
🎯
70%
To Pass
♾️
∞
Time
🔄
∞
Attempts

📝 Instructions

  • Read each question carefully
  • Select the best answer for each question
  • You can retake the quiz as many times as you want
  • Your progress will be shown at the top