Can an HTTP Header Carry Vietnamese? ASCII, obs-text, and Where .NET Draws the Line
Short answer: no, not if you want it to survive every layer. The HTTP specification defines header values in terms of printable ASCII, plus an obsolete branch called obs-text that allows bytes 0x80–0xFF but defines no charset for interpreting them. The consequence is that a Vietnamese string can get through partially: í lives inside Latin-1 and is representable, while ồ is not, because it is U+1ED3 — beyond 0xFF. .NET takes a decisive stance: HttpClient refuses to send any non-ASCII header value. To carry text with diacritics you must encode it, using percent-encoding, RFC 8187 or Base64.
This post branches off a production incident where the header cf-ipcity: Hồ Chí Minh, injected by Cloudflare, made HttpClient throw. There I stopped at the symptom. Here I answer the question that symptom raises: what exactly is an HTTP header allowed to contain?
TL;DR
- An HTTP header value is fundamentally a sequence of bytes, not a Unicode string.
- RFC 9110 allows
0x80–0xFFthrough obs-text, but defines no charset — so no party knows how to decode it reliably. - In practice each side picks a different encoding and you get mojibake.
- Vietnamese gets through partially via Latin-1:
íyes,ồno. - .NET does not gamble:
HttpClientthrows rather than emitting bytes it cannot guarantee. - The correct approaches: percent-encoding, RFC 8187, or Base64.
What a header value actually is, grammatically
Open RFC 9110 and the definition of a header field value reduces to this:
field-value = *field-content
field-content = field-vchar [ 1*( SP / HTAB / field-vchar ) field-vchar ]
field-vchar = VCHAR / obs-text
VCHAR = %x21-7E ; printable ASCII
obs-text = %x80-FF ; non-ASCII bytes, obsolete
Those last two lines are the whole story.
VCHAR is printable ASCII, 0x21 through 0x7E. That part everyone agrees on.
obs-text is bytes from 0x80 to 0xFF. The obs stands for obsolete. It exists in the specification not for you to use, but to describe things that already ended up on the Internet and cannot be removed.
And here is the decisive point: the specification does not say which charset those bytes belong to. Not "it is UTF-8", not "it is Latin-1". Just bytes, with no interpretation attached.
A format that lets you send bytes without telling the receiver what they mean is not a data channel — it is an invitation to misunderstand.
Why í gets through and ồ does not
This is the detail I find most interesting, and it explains why this class of bug shows up intermittently — breaking for some customers and not others.
Take the exact string from the incident and look at its bytes in two encodings:
String : Hồ Chí Minh
UTF-8 : 48 E1 BB 93 20 43 68 C3 AD 20 4D 69 6E 68
Latin-1: 48 3F 20 43 68 ED 20 4D 69 6E 68
^^ ^^
'ồ' → 0x3F 'í' → 0xED
Read character by character:
| Character | Unicode code point | UTF-8 | Latin-1 |
|---|---|---|---|
H | U+0048 | 48 | 48 |
ồ | U+1ED3 | E1 BB 93 | not representable → 3F (?) |
í | U+00ED | C3 AD | ED |
í is U+00ED, comfortably inside 0x00–0xFF, so Latin-1 carries it in a single byte. ồ is U+1ED3 — far beyond 0xFF — so Latin-1 has no room for it and substitutes a question mark.
Vietnamese therefore straddles the boundary: simple accented vowels such as á à í ò ú fall inside Latin-1, while compound forms like ồ ậ ữ ợ ẩ do not. A system passing headers around as Latin-1 will appear to work for Hải and break on Hồ — same code, different data.
That is the worst kind of bug: it depends on content, so it sails past every test you did not deliberately design for it.
What happens when two sides pick different encodings
Suppose one side writes the header as UTF-8 and the other reads it as Latin-1 — the default situation in a great many stacks.
The sender writes ồ as three bytes, E1 BB 93. The receiver reads each byte as Latin-1 and gets three characters: á, », “. The string Hồ becomes Hồ.
This is mojibake, and it is nobody's fault. Both sides did the right thing according to the encoding they chose. The problem is that the specification never made them choose the same one.
What matters is that mojibake is silent. No exception, no warning — just wrong data flowing quietly onward into logs, databases and reports. Compared with .NET throwing an exception outright, silently corrupting data is considerably worse.
Where .NET draws the line
In the incident post I measured this specifically on .NET 9.0.4, and the results are worth repeating because they show .NET taking a clear position:
| Operation | Result |
|---|---|
DefaultRequestHeaders.Add(name, "Hồ Chí Minh") | no exception |
TryAddWithoutValidation(name, "Hồ Chí Minh") | returns true |
SendAsync(request) | throws HttpRequestException: Request headers must contain only ASCII characters. |
| Bytes on the wire | none |
This tells us two things.
First, HttpHeaders validation is not where encoding is checked. It checks structure: control characters, line breaks, per-header format. Whether the value fits in ASCII belongs to the serialisation layer, all the way at send time.
Second, .NET refuses to play the obs-text lottery. It could have chosen to write Latin-1 and let ồ become ?, or to write UTF-8 and let the other side guess. Both lead to silent data corruption. Throwing an exception is the louder choice, but the more honest one.
If .NET's strictness here annoys you, remember that the alternative is not "it works correctly" but "it works incorrectly without telling you".
Three correct ways to put diacritics in a header
When you genuinely need text with diacritics in a header — a filename, a city, a username — the answer is always to encode it into ASCII and decode on the receiving side.
1. Percent-encoding
The most common approach, and the easiest to read in logs:
// Sender
var encoded = Uri.EscapeDataString("Hồ Chí Minh");
// => H%E1%BB%93%20Ch%C3%AD%20Minh
request.Headers.TryAddWithoutValidation("X-City", encoded);
// Receiver
var city = Uri.UnescapeDataString(raw);
The whole result is ASCII, so it survives every layer. Uri.EscapeDataString encodes as UTF-8 and Uri.UnescapeDataString decodes as UTF-8, so both ends agree when both are .NET.
2. RFC 8187 — the standard when you define a new header
RFC 8187 defines a syntax that carries the charset inside the value itself:
X-City*=UTF-8''H%E1%BB%93%20Ch%C3%AD%20Minh
Three parts: the charset name, an optional language tag, then the percent-encoded string. You meet this every day without noticing — it is exactly how filename* works in Content-Disposition when you download a file with a non-ASCII name.
Its advantage over bare percent-encoding: the receiver does not have to guess the charset, because it is declared in place.
3. Base64
Suitable when the data is not plain text, or when you want one uniform rule for every value:
var encoded = Convert.ToBase64String(Encoding.UTF8.GetBytes("Hồ Chí Minh"));
// => SOG7kyBDaMOtIE1pbmg=
The trade-off: logs are no longer human-readable, and the size grows by roughly a third.
Comparison
| Approach | Readable in logs | Self-describing charset | When to use |
|---|---|---|---|
| Percent-encoding | ✅ | ❌ | The default, when both ends share a convention |
| RFC 8187 | ✅ | ✅ | When you define a new header for several consumers |
| Base64 |
