Yes! You can upload a corpus of text that the service will learn new words (and their context) from and/or you can tell it it specific words and their pronunciation. No audio samples are needed, the customization works on the existing language models.
More details: https://www.ibm.com/watson/developercloud/doc/speech-to-text...