What did Sentence-Bert return here?
What did Sentence-Bert return here?
Loading saved threads...
user900476 · External communityPost link
External question — Data Science Stack Exchange
Author: user900476
Original post: https://datascience.stackexchange.com/questions/112402
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I used sentence bert to embed sentences from this tutorial
https://www.sbert.net/docs/pretrained_models.html
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer('all-mpnet-base-v2')
This is the event triples
t
I forgot to concat into sentences,
[('U.S. stock index futures', 'points to', 'start'),
('U.S. stock index futures', 'points to', 'higher start')]
model.encode(t)
returns a 2d array of shape (2,768), with two idential 768-dimension vectors, and its value is different from both
model.encode('U.S. stock index futures')
and
model.encode('U.S. stock index futures points to start')
. What could possibly have it returned?
It is the same situation for other models on huggingface such as
https://huggingface.co/sentence-transformers/stsb-distilbert-base
Quote
Report
Pushpam Punjabi · External communityPost link
External answer — Data Science Stack Exchange
Author: Pushpam Punjabi
Original post: https://datascience.stackexchange.com/a/112405
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
There are two valid inputs to MPNet's tokenizer:
Union[TextInputSequence, Tuple[InputSequence, InputSequence]]
When you give a list of tuples as an input, from each tuple only the first two sentences, i.e. "U.S. stock index futures" and "points to" are used to encode, similar to BERT's next sentence prediction pre-training task.
This is tokenized and converted to input_ids with special tokens as follows:
['<s>', 'u', '.', 's', '.', 'stock', 'index', 'futures', '</s>', '</s>', 'points', 'to', '</s>']
<s> indicates start of sentence and </s> indicates end of sentence. I'm not sure about the </s> before 'points', but this is the output from the model's tokenizer. As such NLU models can work with one or two sentences together, these special tokens are added to identify how many sentences are in the model and they help to separate the sentences.
Thus, you get two same vectors from
model.encode(t)
If you want the same vector without using a tuple, the encode function should contain the following input string:
"U.S. stock index futures </s> </s> points to"
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Pushpam Punjabi Source score (net votes, not local likes): 2 Original post: https://datascience.stackexchange.com/a/112405 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. There are two valid inputs to MPNet's tokenizer: Union[TextInputSequence, Tuple[InputSequence, InputSequence]] When you give a list of tuples as an input, from each tuple only the first two sentences, i.e. "U.S. stock index futures" and "points to" are used to encode, similar to BERT's next sentence prediction pre-training task. This is tokenized and converted to input_ids with special tokens as follows: ['<s>', 'u', '.', 's', '.', 'stock', 'index', 'futures', '</s>', '</s>', 'points', 'to', '</s>'] <s> indicates start of sentence and </s> indicates end of sentence. I'm not sure about the </s> before 'points', but this is the output from the model's tokenizer. As such NLU models can work with one or two sentences together, these special tokens are added to identify how many sentences are in the model and they help to separate the sentences. Thus, you get two same vectors from model.encode(t) If you want the same vector without using a tuple, the encode function should contain the following input string: "U.S. stock index futures </s> </s> points to"
Checking account access…