Watson Discovery uses stemming to identify valid matches. Based on our experience it appears that stemming is applied in both Natural Language Queries and also Discovery Language Queries. It seems to also enforced in phrase searching. This is useful sometimes but there are other cases where this approach makes it very difficult to find relevant matches. For example, in one of our collections we have several documents with the term "DCS" and several with the term "DC". These are not related terms and there is no reason for our users to get both in a single query. However, if you search for either DCS or DC, you will get both and there is no way to filter our the undesired matches because they seem to be interpreted by WDS as the exact same term.
We think WDS should not enforce stemming when using phrase searching, as this type of search is primarily used to specify exact matches such as titles or excerpts of a document. Alternatively, we would like to have some control over when stemming should be applied or what words or terms should not be stemmed.
|Who would benefit from this IDEA?||"As a customer I would like to be able to do searches that only retrieve exact matches|
NOTICE TO EU RESIDENTS: per EU Data Protection Policy, if you wish to remove your personal information from the IBM ideas portal, please login to the ideas portal using your previously registered information then change your email to "firstname.lastname@example.org" and first name to "anonymous" and last name to "anonymous". This will ensure that IBM will not send any emails to you about all idea submissions