std::codecvt_utf8_utf16
cppreference.com
[] |
Google. . , . , . |
<metanoindex/>
<tbody> </tbody> template< class Elem, unsigned long Maxcode = 0x10ffff, std::codecvt_mode Mode = (std::codecvt_mode)0 > class codecvt_utf8_utf16 : public std::codecvt<Elem, char, std::mbstate_t>; |
||
std::codecvt_utf8_utf16 std::codecvt , UTF-8 UTF-16
Elem 32... - , UTF-16 32- . codecvt UTF-8 , , .:
std::codecvt_utf8_utf16 is a std::codecvt facet which encapsulates conversion between a UTF-8 encoded byte string and UTF-16 encoded character string. If
Elem is a 32-bit type, one UTF-16 codepoint will be stored in each 32-bit character of the output sequence. This codecvt facet can be used to read and write UTF-8 files, both text and binary.| Elem | ||
| Maxcode | Elem, | |
| Mode |
std::codecvt
Member types
intern_type
|
internT
|
extern_type
|
externT
|
state_type
|
stateT
|
Member objects
| Type | |
id ()
|
std::locale::id |
Member functions
do_out (public - std::codecvt)
| |
do_in (public - std::codecvt)
| |
do_unshift (public - std::codecvt)
| |
do_encoding (public - std::codecvt)
| |
do_always_noconv (public - std::codecvt)
| |
do_length (public - std::codecvt)
| |
do_max_length (public - std::codecvt)
|
Protected member functions
[virtual] |
internT externT, , (virtual protected std::codecvt -)
|
[virtual] |
externT internT, , (virtual protected std::codecvt -)
|
[virtual] |
externT : generates the termination character sequence of externT characters for incomplete conversion (virtual protected std::codecvt -)
|
[virtual] |
externT , internT , : returns the number of externT characters necessary to produce one internT character, if constant (virtual protected std::codecvt -)
|
[virtual] |
, (virtual protected std::codecvt -)
|
[virtual] |
externT , internT : calculates the length of the externT string that would be consumed by conversion into given internT buffer (virtual protected std::codecvt -)
|
[virtual] |
externT , internT : returns the maximum number of externT characters that could be converted into a single internT character (virtual protected std::codecvt -)
|
UTF -8 UTF-16 32-
wchar_t
:
The following example demonstrates reading a UTF-8 file into a UTF-16 string on a system with 32-bit
wchar_t
#include <fstream>
#include <iostream>
#include <string>
#include <locale>
#include <codecvt>
int main()
{
std::ofstream("text.txt") << u8"z\u6c34\U0001d10b";
std::wifstream file1("text.txt");
file1.imbue(std::locale("en_US.UTF8"));
std::cout << "Normal read from file (using default UTF-8/UTF-32 codecvt)\n";
for(wchar_t c; file1 >> c; )
std::cout << std::hex << std::showbase << c << '\n';
std::wifstream file2("text.txt");
file2.imbue(std::locale(file2.getloc(), new std::codecvt_utf8_utf16<wchar_t>));
std::cout << "UTF-16 read from the same file (using codecvt_utf8_utf16)\n";
for(wchar_t c; file2 >> c; )
std::cout << std::hex << std::showbase << c << '\n';
}
:
Normal read from file (using default UTF-8/UTF-32 codecvt)
0x7a
0x6c34
0x1d10b
UTF-16 read from the same file (using codecvt_utf8_utf16)
0x7a
0x6c34
0xd834
0xdd0b
.
| Character conversions |
narrow multibyte (char) |
UTF-8 (char) |
UTF-16 (char16_t) |
|---|---|---|---|
| UTF-16 | mbrtoc16 / c16rtomb | codecvt<char16_t, char, mbstate_t> codecvt_utf8_utf16<char16_t> codecvt_utf8_utf16<char32_t> codecvt_utf8_utf16<wchar_t> |
/ |
| UCS2 | codecvt_utf8<char16_t> | codecvt_utf16<char16_t> | |
| UTF-32/UCS4 (char32_t) |
mbrtoc32 / c32rtomb | codecvt<char32_t, char, mbstate_t> codecvt_utf8<char32_t> |
codecvt_utf16<char32_t> |
| UCS2/UCS4 (wchar_t) |
codecvt_utf8<wchar_t> | codecvt_utf16<wchar_t> | |
| wide (wchar_t) |
codecvt<wchar_t, char, mbstate_t> mbsrtowcs / wcsrtombs |
| , UTF-8, UTF-16, UTF-32 ( ) | |
(C++11)( C++17) |
codecvt () |
(C++11)( C++17) |
UTF-8 UCS-2/UCS-4 ( ) |
(C++11)( C++17) |
UTF-16 UCS-2/UCS-4 ( ) |