I'm creating this separate thread for paragraph based segmentation issues I mentioned elsewhere:
https://community.sdl.com/solutions/language/translationproductivity/f/90/p/13299/46835#46835
(additional info in https://community.sdl.com/product-groups/translationproductivity/f/90/p/13299/48428#48428)
I'm attaching test files here:
5383.testfiles.zip
ZIP contains:
IAD_Pseudo-HTML.sdlftsettings - file type definition
test_en-US_de-DE.sdltm - testing TM with the additional Line break segmentation rule
test_English.txt.html - source file to test with
test_English.txt - original client's source format for reference only (it's very weird hybrid of plain text and HTML content with some custom non HTML-compliant tags and some "Excel formulas-like code"... normally it contains more of that 'garbage', but I deleted it as it's not relevant now)
Now, already in the "working" version I see some weirdness in the LineBreak rule - the only rule which works is this weird one:
.?[\r\n][\r\n]
The logical one - .?[\r\n]+ - does NOT work for unknown reason :-O
But at least something works...
In Studio 2015 SR3 (12.3.5262.0) and 2017 CU5 (14.0.5889.5) I get this - all properly segmented as expected...
In Studio 2015 SR3 CU10 (12.3.5281.10) and 2017 SR1 CU7 (14.1.6329.7) I get this - segments broken at totally weird places in a middle of a word... but at the same time NOT broken at the line break (sometimes :-O)
It must have something to do with the inline locked content, since plaintext lines further down the test file are segmented correctly.
And more weirdly, I am not able to find regex which would make it work in these new Studio updates!
Neither the basic .?[\r\n]+ regex, nor the [\w\p{P}][\r\n]+ regex according this knowledge base article work - the basic regex produces exactly the same weirdness as above, the regex from the KB article gets better, but fails on lines with trailing spaces (and adding \s* anywhere in the regex yields again exactly the same weird result as above):
So, my conclusion is that there is (still) something broken in the segmentation rules.
Translate

