You are currently working on UAT 

PO (portable object) file segmentation on a sentence basis?

Hi,

I need help for PO file segmentation.

As the PO format is of a bilingual type, Trados Studio each pair of "msgid" and "msgstr" as a single segment.

Sometimes, however, this segmentation rule is annoying especially when a single pair contains multiple sentences.

Is there any way to segment a PO file on a sentence basis, regardless of how many sentences are contained in each "msgid" field?

I think it won't be technically difficult as the bilingual Excel type in Studio already supports the same segmentation rule.

  • Hi

    The bilingual Excel file type you mention does not segment content using the set segmentation rules if there is any content in the target language column. This is most likely for the same reason:

    As soon as you segment both languages, you'd need to have some form of an alignment process.

    Example:

    Source: This is a rather simple example. But in this context, it may suffice.

    Target: Dieses eher einfache Beispiel dürfte unter diesem Umständen ausreichen.

    Or, alternative translation: Unter diesen Umständen mag Folgendes ausreichen: Ein eher einfaches Beispiel.

    The translations themselves might be deficient, but I hope the point is clear.

    I would hope for SDL to develop a routine for all bilingual file types which does the segmentation of the source text and then opens the aligner to let the translator align existing target content. I don't think that would be very hard to do because they have all the building blocks for that, but such a thing does not yet exist, at least not in the SDL universe.

    Daniel

  • Thank you for your reply.

    In the case of the bilingual Excel file, if the target cells are blank, Studio conforms to the segmentation rule.

    (Only when the target is already filled, the file is segmented as you said.)

    And that's enough for me because in most cases, the reason why I use Trados Studio is to make a new translation from scratch.

    But PO files are always segmented based on its native TUs....

  • It's not only Studio that will do this.  If you open such a file in POEdit for example you'll see something like this:

    POEdit interface showing English source text with repeated 'Unknown system error' messages and their Icelandic translations with variations of 'Obekkt kerfisvilla'.

    The problem for Studio (and I assume for others too) is that we have no idea how to segment the file correctly 100% of the time unless it's like this.  If you had three sentences in the source and only two in the target what would you do?  I know a guess might work, but that's not very technical, so the segmentation controls with these sort of files that provide a structural environment in the file itself is to treat them as paragraphs.

    Paul Filkin | RWS

    Design your own training!
    You've done the courses and still need to go a little further, or still not clear? 
    Tell us what you need in our Community Solutions Hub

    emoji


    Generated Image Alt-Text
    [edited by: Trados AI at 4:14 PM (GMT 0) on 28 Feb 2024]