Reference

Specific Character Set

The tag (0008,0005) that names the character sets a file uses for text, such as Japanese, Korean or Chinese.

Browser: available nowmacOS app: 1.4.0 or newer

Specific Character Set (0008,0005) tells a reader how to turn the bytes of a text value (VR SH, LO, ST, LT, UT, PN or UC) into characters. With no value, text is plain ASCII. ISO_IR 100 is Latin-1, and ISO_IR 192 is UTF-8. A sequence item can declare its own set, which applies inside that item only.

ISO 2022 code extensions

Japanese, Korean and Chinese files often use ISO 2022 code extensions. The tag then holds several terms, for example \ISO 2022 IR 87 for ASCII plus JIS X 0208 kanji. Inside a value, escape sequences switch between the named sets. A Japanese patient name often holds three forms of the same name this way: Yamada^Tarou=山田^太郎=やまだ^たろう.

dcmage opens these files and decodes each value by the sets its escape sequences switch to: ISO 2022 IR 6, 13, 58, 87, 100, 101, 109, 110, 126, 127, 138, 144, 148, 149, 159, 166 and 203. That covers Japanese (JIS X 0201, JIS X 0208, JIS X 0212), Korean (KS X 1001), simplified Chinese (GB 2312), Thai (TIS 620) and the ISO 8859 sets.

What changes on export

dcmage writes every text value as UTF-8 and sets Specific Character Set to ISO_IR 192. The characters stay the same, but the bytes and the tag value differ from the source file. A receiving system has to accept UTF-8.

Text dcmage cannot decode

dcmage cannot decode a run of bytes after an escape sequence for a set outside the list above, such as JIS X 0213, or bytes that are not valid in the set in force. The file still opens. The tag tree shows the part it could not read as � and marks the row undecodable.

Export is refused while any value still holds such a part. The notice names each element, because writing it would replace the original characters with � for good. Edit or delete the value, then export again.