{"id":174,"date":"2007-11-14T16:36:04","date_gmt":"2007-11-14T23:36:04","guid":{"rendered":"http:\/\/www.chesnok.com\/daily\/2007\/11\/14\/automatic-character-set-conversion-in-postgresql\/"},"modified":"2012-03-26T02:57:10","modified_gmt":"2012-03-26T10:57:10","slug":"automatic-character-set-conversion-in-postgresql","status":"publish","type":"post","link":"https:\/\/www.chesnok.com\/daily\/2007\/11\/14\/automatic-character-set-conversion-in-postgresql\/","title":{"rendered":"automatic character set conversion in postgresql"},"content":{"rendered":"<p>Today, I encountered a few goofy characters in the data I am migrating from one ERP system to another. For example, &#8220;\u00c2\u00a2&#8221; isn&#8217;t represented the same way in UTF-8 as LATIN1 character sets. In UTF-8, the hex representation for &#8220;\u00c2\u00a2&#8221; is <code>c2 a2<\/code>, but in LATIN1 it is  <code>a2<\/code>. <\/p>\n<p>I started looking for an easy Perl way to translate everything into UTF-8 on the client side, when I discovered that PostgreSQL offers <a href=\"http:\/\/www.postgresql.org\/docs\/8.2\/static\/multibyte.html#AEN24142\">automatic client-to-server character set conversions<\/a>.  All I have to do is specify what my client character set is. <\/p>\n<p>Here&#8217;s how you can do it with an SQL command: <\/p>\n<blockquote><p>\n<code>SET CLIENT_ENCODING TO 'LATIN1';<br \/>\n<\/code>\n<\/p><\/blockquote>\n<p>Substitute your character set for &#8220;LATIN1&#8221;. <\/p>\n<p>Lucky for me, my database is set to <code>UTF8<\/code>, and in that case, all supported encodings on my clients will be automatically converted to UTF-8 &#8212; as long as I specify which encoding I&#8217;m using.<\/p>\n<p>The support for UTF-8 (formerly called UNICODE in the docs) in PostgreSQL has been around since version 7.1 (early 2000), and in version 8.1 the conversion support for UTF-8 was expanded to all known character sets.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Today, I encountered a few goofy characters in the data I am migrating from one ERP system to another. For example, &#8220;\u00c2\u00a2&#8221; isn&#8217;t represented the same way in UTF-8 as LATIN1 character sets. In UTF-8, the hex representation for &#8220;\u00c2\u00a2&#8221; &hellip; <a href=\"https:\/\/www.chesnok.com\/daily\/2007\/11\/14\/automatic-character-set-conversion-in-postgresql\/\">Continue reading &rarr;<\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[97,9],"tags":[647],"class_list":["post-174","post","type-post","status-publish","format-standard","hentry","category-postgres","category-postgresql","tag-postgres"],"_links":{"self":[{"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/posts\/174","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/comments?post=174"}],"version-history":[{"count":1,"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/posts\/174\/revisions"}],"predecessor-version":[{"id":4057,"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/posts\/174\/revisions\/4057"}],"wp:attachment":[{"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/media?parent=174"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/categories?post=174"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.chesnok.com\/daily\/wp-json\/wp\/v2\/tags?post=174"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}