You are currently working on UAT 

4.75 GB SDL TM - now running in Studio 2017 SR1

I have been "stashing" more than 15 years of TUs in my "personal" TM. It has now grown to a rather large size about 5GB and more than 700000 TUs.

Do you see that as a problem?

I know Access databases usually don't function well if over 2 GB. - Is this an issue for the .sdltm file? - and if yes, what do you suggest I do.

My plan so far:

1) I am going to start looking for duplicates to reduce the size, but that is such a tedious process using just the Detail View in the TM interface and looking for duplicates there - Numbers, even part numbers (in letter form) and tags apparently aren't recognized as differences.

2) I know I can export to .tmx and then re-import into a new TM.

If I do that, what do you recommend for settings to strip out duplicates, or TUs that maybe differ only by tags. - Or is that even a process that may work?

Are there any good tutorials out there?

Thanks you all,

Susanna

Parents
  • Hi  

    Nagoya is mostly correct, but just to clarify.

    Susanna R Miles said:
    I know Access databases usually don't function well if over 2 GB. - Is this an issue for the .sdltm file? - and if yes, what do you suggest I do.

    Studio uses SQLite for Translation Memories and these in theory can handle up to 140 TB, so 5GB isn't an issue.  You can read more if you're interetsed here:

    https://www.sqlite.org/limits.html

    Susanna R Miles said:
    Numbers, even part numbers (in letter form) and tags apparently aren't recognized as differences.

    This should reduce the size of your TM.  Recognized numbers don't need to be duplicated since Studio sees then all as one and "should" be able to autosubstitute the correct value without storing every single number you ever translate.

    Susanna R Miles said:

    2) I know I can export to .tmx and then re-import into a new TM.

    If I do that, what do you recommend for settings to strip out duplicates, or TUs that maybe differ only by tags

    The default settings would strip out true dupicates.  If the difference was only tags then you should probably leave them in there as these are not duplicates.  If you remove the TUs with tags then you could get a 99% match instead of a 100% for a document you have translated before with tags.

    You could also have a play with this application which has a feature for duplicate removal:

    https://appstore.sdl.com/language/app/sdl-translation-memory-management-utility/131/

    But make sure you back up your TM before you start... I'd recommend you make a back up to TMX as well and do these backups regularly if you don't already.

    Paul Filkin | RWS

    Design your own training!
    You've done the courses and still need to go a little further, or still not clear? 
    Tell us what you need in our Community Solutions Hub

Reply
  • Hi  

    Nagoya is mostly correct, but just to clarify.

    Susanna R Miles said:
    I know Access databases usually don't function well if over 2 GB. - Is this an issue for the .sdltm file? - and if yes, what do you suggest I do.

    Studio uses SQLite for Translation Memories and these in theory can handle up to 140 TB, so 5GB isn't an issue.  You can read more if you're interetsed here:

    https://www.sqlite.org/limits.html

    Susanna R Miles said:
    Numbers, even part numbers (in letter form) and tags apparently aren't recognized as differences.

    This should reduce the size of your TM.  Recognized numbers don't need to be duplicated since Studio sees then all as one and "should" be able to autosubstitute the correct value without storing every single number you ever translate.

    Susanna R Miles said:

    2) I know I can export to .tmx and then re-import into a new TM.

    If I do that, what do you recommend for settings to strip out duplicates, or TUs that maybe differ only by tags

    The default settings would strip out true dupicates.  If the difference was only tags then you should probably leave them in there as these are not duplicates.  If you remove the TUs with tags then you could get a 99% match instead of a 100% for a document you have translated before with tags.

    You could also have a play with this application which has a feature for duplicate removal:

    https://appstore.sdl.com/language/app/sdl-translation-memory-management-utility/131/

    But make sure you back up your TM before you start... I'd recommend you make a back up to TMX as well and do these backups regularly if you don't already.

    Paul Filkin | RWS

    Design your own training!
    You've done the courses and still need to go a little further, or still not clear? 
    Tell us what you need in our Community Solutions Hub

Children
No Data