{"id":8432,"date":"2018-01-03T14:22:50","date_gmt":"2018-01-03T13:22:50","guid":{"rendered":"http:\/\/www.cpsl.com\/2018\/01\/03\/lsp-perspective-applying-human-touch-mt-qualitative-feedback-mt-evaluation-part-2\/"},"modified":"2025-04-02T14:25:19","modified_gmt":"2025-04-02T14:25:19","slug":"lsp-perspective-applying-human-touch-mt-qualitative-feedback-mt-evaluation-part-2","status":"publish","type":"post","link":"https:\/\/www.cpsl.com\/en-GB\/lsp-perspective-applying-human-touch-mt-qualitative-feedback-mt-evaluation-part-2\/","title":{"rendered":"LSP Perspective: Applying the Human Touch to MT (Part 2)"},"content":{"rendered":"<h2>In all the discussions about <a href=\"http:\/\/www.cpsl.com\/en-GB\/services\/ai-smart-machine-translation\/\" target=\"_blank\" rel=\"noopener noreferrer\">Machine Translation<\/a>, we do not often hear much about post-editors and what could be done to enhance and improve the task of PEMT, which is often viewed negatively.\u00a0 Our expert<strong>\u00a0<\/strong>provides useful insights into her direct experience of improving the experience for post-editors doing this kind of work.<\/h2>\n<p><em>(PART 2)<\/em>&#8230; More than asking for severity and repetitiveness,<strong> what I really want to know is what I call \u2018annoyance level,\u2019 i.e. what made the post-editing job too boring, tedious or time-consuming \u2013 in short, a task that could lead the post-editor to decline a similar job in the future. These are variables that quantitative metrics cannot provide.<\/strong> Automated metrics cannot provide any insight on how to prioritise error fixing, either by error severity level or by \u2018annoyance level.\u2019 Important errors can go unnoticed in a long list of issues, and thus never be fixed. I have managed several MT-based projects where the edit distance was acceptable (&lt; 30%) and the post-editors\u2019 overall experience, to my surprise was still unpleasant. In such cases, the post-editor came back to me saying that certain types of errors were so unacceptable for them that they didn\u2019t want to post-edit again. Sometimes this opinion was related to severity and other times to perception, i.e. errors a human would never make. In these cases, the feedback form helped detect the errors and turned a previously bad experience into an acceptable job.<\/p>\n<p>It is worth noting that one cannot rely on one single post- editor&#8217;s feedback. The acceptance threshold can vary quite a lot from one person to another, and post-editing skills are also different. Thus, the most reasonable approach is to collect feedback from several post-editors, compare their comments and use them as a complement to the automatic metrics. We must definitely make an effort to include the post-editors\u2019 comments as a variable when evaluating MT quality, to prioritise certain errors when optimising the engines. If we have a team of translators whom we trust, then we should also trust them when they comment on the raw MT results. Personally, I always try my best to send machine-translated files that are in good shape so that the post-editing experience is acceptable. In this way, I can keep my preferred translators (recycled as post-editors) happy and on board, willing to accept more jobs in the future. This can make a significant difference not only in their experience but also in the quality of the final project.<\/p>\n<div class=\"mceTemp\"><\/div>\n<h3><strong>5 Tips for Successfully Integrating Qualitative Feedback into your MT Evaluation Workflow<\/strong><\/h3>\n<ul>\n<li><strong>Devise a tool and a workflow for collecting feedback from the post-editors.<\/strong><\/li>\n<\/ul>\n<p>It doesn\u2019t have to be a sophisticated tool and the post-editors shouldn\u2019t have to fill in huge Excel files with all changes and comments. It\u2019s enough to collect the most awkward errors; those they wouldn\u2019t want to fix over and over again. However, if you don\u2019t have the time to read and process all this information, a short informal conversation on the phone from time to time can also be of help and give you valuable feedback about how the system is working.<\/p>\n<ul>\n<li><strong>Agree to fair compensation<\/strong><\/li>\n<\/ul>\n<p>Much has been said about this. My advice would be to rely on the automatic metrics but to include the post- editor&#8217;s feedback in your decision. Therefore, I usually offer hourly rates when language combinations are new and the effort is harder, and per word rates when the MT systems are established and have stable edit distances. When using hourly rates, you can ask your team to use time-tracking apps in their CAT tools or ask them to report the real hours spent. To avoid last-minute surprises, for full PE it is advisable to indicate a maximum number of hours based on the expected PE speed, and ask them to inform you of any deviation, whereas for light post-editing you may want to indicate a minimum number of hours to make sure the linguists are not leaving anything unchecked.<\/p>\n<ul>\n<li><strong>Never promise the moon<\/strong><\/li>\n<\/ul>\n<p>If you are running a test, tell your team. Be honest about the expected quality and always explain the reason why you are using MT (cost, deadline\u2026).<\/p>\n<ul>\n<li><strong>Don\u2019t force anyone to become a post-editor<\/strong><\/li>\n<\/ul>\n<p>I have seen very good translators becoming terrible post-editors; either they change too many things or too few, or simply cannot accept that they are reviewing a translation done by a machine. I have also seen bad translators become very good post-editors. Sometimes a quick chat on the phone can be enough to check if they are reluctant to use MT per se, or if the system really needs further improvement before the next round.<\/p>\n<ul>\n<li><strong>Listen, listen, listen<\/strong><\/li>\n<\/ul>\n<p>We PMs tend to provide the translators with a lot of instructions and reference material and make heavy use of email. Sometimes, however, it\u2019s worth it to arrange short calls and listen to the post-editors\u2019 opinion of the raw MT. For long-term projects or stable MT-based language combinations, it is also advisable to arrange regular group calls with the post-editors, either by language or by domain.<\/p>\n<h3><strong>And\u2026 What About NMT Evaluation?<\/strong><\/h3>\n<p>According to several studies on NMT, it seems that the errors produced by these systems are harder to detect than those produced by RBMT and SMT, because they occur at the semantic level (i.e. meaning). NMT takes context into account and the resulting text flows naturally; we no longer see the syntactically awkward sentences we are used to with SMT. But the usual errors are mistranslations, and mistranslations can only be detected by post-editors, i.e. by people. In most NMT tests done so far, BLEU scores were low, while human evaluators considered that the raw MT output was acceptable, which means that with NMT we cannot trust BLEU alone. Both source and target text have to be read and assessed in order to decide if the raw MT is acceptable or not; human evaluators have to be involved. With NMT, human assessment is clearly even more important, so while the translation industry works on a valid approach for evaluating NMT, it seems that qualitative information will be required to properly assess the results of such systems.<\/p>\n<p>&nbsp;<\/p>\n<div id=\"attachment_6191\" style=\"width: 563px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/www.cpsl.com\/services\/post-editing\/\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-6191\" class=\" wp-image-6191\" src=\"http:\/\/www.cpsl.com\/wp-content\/uploads\/2017\/11\/Post-Editor-Feeback-Form.png\" alt=\"CPSL - Post Editor Feeback Form\" width=\"553\" height=\"257\" \/><\/a><p id=\"caption-attachment-6191\" class=\"wp-caption-text\"><strong>Post Editor Feedback Form<\/strong><\/p><\/div>\n<p>CPSL has been the driving force behind the new <a href=\"https:\/\/atc.org.uk\/iso-certification-service\/standards\/\" target=\"_blank\" rel=\"noopener noreferrer\"><strong>ISO 18587 (2017)<\/strong> standard.<\/a>\u00a0ISO 18587 (2017) regulates the post-editing of content processed by machine translation systems, and also establishes the competences and qualifications that post-editors must have. The standard is intended for use by post-editors, translation service providers and their clients.<\/p>\n<h4 class=\"shortened_begin2\"><strong><a href=\"https:\/\/www.youtube.com\/watch?v=G8GMEcSETmE&amp;feature=youtu.be\" target=\"_blank\" rel=\"noopener noreferrer\">Click here<\/a> to see some of CPSL best MT practices to familiarise yourself with this genre.<br \/>\n<\/strong><\/h4>\n","protected":false},"excerpt":{"rendered":"<p>In all the discussions about Machine Translation, we do not often hear much about post-editors and what could be done to enhance and improve the task of PEMT, which is often viewed negatively.\u00a0 Our expert\u00a0provides useful insights into her direct experience of improving the experience for post-editors doing this kind of work. (PART 2)&#8230; More [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":6217,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"aside","meta":{"inline_featured_image":false,"footnotes":""},"categories":[127,58],"tags":[],"class_list":["post-8432","post","type-post","status-publish","format-aside","has-post-thumbnail","hentry","category-post-editing-gb","category-translation-gb","post_format-post-format-aside"],"_links":{"self":[{"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/posts\/8432","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/comments?post=8432"}],"version-history":[{"count":17,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/posts\/8432\/revisions"}],"predecessor-version":[{"id":40174,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/posts\/8432\/revisions\/40174"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/media\/6217"}],"wp:attachment":[{"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/media?parent=8432"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/categories?post=8432"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.cpsl.com\/en-GB\/wp-json\/wp\/v2\/tags?post=8432"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}