Sign in

Aaron Tay

@aarontay.bsky.social
3.5K followers 345 following 4.8K posts

Currently on organizing committee of FORCE2026, Singapore, 3-5 June 2026. I'm librarian + blogger from Singapore Management University. Social media, bibliometrics, analytics, academic discovery tech.

PostsRepliesMedia
Aaron Tay @aarontay.bsky.social · 08/10/2026
Yes. I was wondering about that too.
000
Aaron Tay @aarontay.bsky.social · 07/10/2026
I can see how this might be used together with COUNTER for usage tracking and might have some privacy implications. It also feels like such a system will have other implications..hmm..
100
Aaron Tay @aarontay.bsky.social · 07/10/2026
NISO wants to define a minimum provenance metadata package that travels with scholarly content when an AI system retrieves & uses it. My guess is In context of al search using naive RAG, think of it as standardising metadata around the grounding/context that is useed by the ai so that is traceable?
100
Aaron Tay @aarontay.bsky.social · 07/10/2026
3 main projects. 1. how AI use of scholarly content is measured and reported(COUNTER) 2.how provenance and Attribution remain linked to the AI agent and hopefully output (NISO) and 3. AI agents discover and access scholarly content (Ithaka S+R). They all look fascinating but 2 especially intrigue me
210
Aaron Tay @aarontay.bsky.social · 07/10/2026
Trusted Retrieval & Attribution for Content Ecosystems (TRACE) brings today various stakeholders and provides "a coordinating forum that helps align and amplify complementary efforts already underway across the ecosystem."
110
Aaron Tay @aarontay.bsky.social · 07/10/2026
First the intro is drawing a historical parallel to the transition from print to online. Just as new infrastructure needed to be bulit then doi, COUNTER, KBART, we need new infrastructure to handle a world where AI agents are searching, accessing full text, summarising, citing on our behalf
120
Aaron Tay @aarontay.bsky.social · 07/10/2026
That said, the article was written in a quite technical dense way and I had to read in twice then asked ChatGPT for help to confirm my understanding. I suspect many librarians will struggle similarly
100
Aaron Tay @aarontay.bsky.social · 07/10/2026
Fascinating .. so many interesting questions this opens up.
110
Aaron Tay @aarontay.bsky.social · 07/10/2026
You might be surprised to know this but I'm uncomfortable with the idea of the library spending money on buying books that are ai generated. It's particularly bad if all you get is some corporate author...and not named individual authors /1
040
Aaron Tay @aarontay.bsky.social · 07/10/2026
Yeah typo. Humans do that
110
Aaron Tay @aarontay.bsky.social · 07/10/2026
Yes. There's very little detail except gen ai is used to "help"
100
Aaron Tay @aarontay.bsky.social · 06/10/2026
[blogged] Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-...
aarontay.substack.com
Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening?
TL;DR: I tested Jev, a cheap zero-shot decision model, for title/abstract screening in systematic reviews.
061
Aaron Tay @aarontay.bsky.social · 06/10/2026
Nice post. My hunch though is it is making a subset of people more ambitious and lure them into understanding and doing more. Then there are those who just accept what is given
040
Aaron Tay @aarontay.bsky.social · 05/10/2026
These embeddings we use are now based on transformer models a type of neural net architecture. Its has black box has how your brain neurons work ;p.
000
Aaron Tay @aarontay.bsky.social · 05/10/2026
My sense is librarians will be very frustrated by non boolean methods. Even lexical non boolean (yes not all lexical methods are boolean ). It's almost like a dark art. Why does this method work? Often its not clear except someone did a bunch of tests and results was better usually!
020
Aaron Tay @aarontay.bsky.social · 05/10/2026
I think knowing only boolean for retrieval will result in very bad intutions on how embeddings work. Its a different paradigm. There are so many common misconceptions by librarians on how retrieval works. I know cos I had them!
110
Aaron Tay @aarontay.bsky.social · 05/10/2026
With semantic search involving embeddings is really hard to tell what types of expansions helps. This isn't like boolean where you can quite easily inspect results and reason why. It's complicated by fact there's not one type of embeddings , and often are fine tuned with examples so they are diff
210
Aaron Tay @aarontay.bsky.social · 05/10/2026
I acutally agree. The only one I know of who has done this is web of science smart search. I have been asking for it for a while and they the only ones who have done it. You can filter to show just semantic or boolean. And in the combined list, they label semantic only hits
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
is the don't understand but think they know better people that I think is the problem...
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
I am also going to guess most people will not realise for such semantic expansion, the expansion doesnt even need to be CORRECT to improve results! The rough form of the answer typically helps push the query embedding closer. I can imagine people not getting this and insisting the expansion is bad.
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
think about it.. why are vendors showing BOOLEAN expansions which is a type of query expansion but hiding the query expansions for semantic search? My guess? people don't understand it, and complain. Both are query expansion, but they work differently. Boolean is known and understood
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
Let just stick to our lane, and ask questions about the index/sources it covers and leave the retrieval algos to the Info retrieval folks. But i notice librarians love to talk about *bias* , and i cant see how you can really talk about it without really understanding info retrieval at some level
110
Aaron Tay @aarontay.bsky.social · 05/10/2026
Part of me thinks besides evidence synthesis cases perhaps, who cares what is going on under the hood. I can know exactly every step of the retreival pipeline and it still doesnt help if the results are bad. Or vice versa. Do we librarians really need to know? Most librarians arent trained anyway/10
200
Aaron Tay @aarontay.bsky.social · 05/10/2026
Thats a good point though people might argue back they know exactly the keywords they enter has to be there.. but HYDE is literally just making things up and often can generate even clearly false or wrong statements (but still works)!
010
Aaron Tay @aarontay.bsky.social · 05/10/2026
so some of them when they dont understand how a search works because they lack sufficient information retrieval training will get very unhappy and say the search is broken etc or insist they know better & vendor implemented wrong. For a vendor a smarter choice would be to hide all these details. /9
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
In fact, if I were the vendor I would be wondering how the heck can I even show this to users? Here's a uncomfortable truth. I will get a lot of flake for this, but I know librarians and i have seen this happen over my career but librarians have an image of themselves as being expert in search/8
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
But think about it a doc that might be the answer you need, when converted to embedding is more likely to be more similar to emebdding of the actual doc with the answer! Also LLMs have been trained on a lot.they may generate a "fake doc" that actually has correct answer but thats not necessary /7
210
Aaron Tay @aarontay.bsky.social · 05/10/2026
yes i was using the API via POST method. didnt even try the web form.
000
Aaron Tay @aarontay.bsky.social · 05/10/2026
to the eyes of most people including librarians this query expansion looks INSANE. Hypothetical Document Embeddings (HyDE) sounds crazy the first time you encounter the idea. You literally ask the LLM to generate a fake doc that might answer the question and embed that rather than the real query! /6
110
Aaron Tay @aarontay.bsky.social · 05/10/2026
Here's what they are actually doing. They take the original query + short synthetic passage describing what a relevant document might talk about and use that to create the query embedding. If you into IR, you recognise this as close in spirit to HyDE / query2doc-style expansion. But otherwise.. /5
210
Aaron Tay @aarontay.bsky.social · 05/10/2026
So why dont vendors just display the query they use to convert to embeddings? Sure, even with that it is still very hard to understand why a certain result has a cosine similarity of 0.4 with the query embedding but at least that's more transparency right? Those vendors are evil for hiding info! /4
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
Below is a mock up of what the UI showed. But this is a very typical hybrid search of lexical and semantic (dense embedding matching). Most systems now do tell you the exact boolean string they used but are silent about the semantic part. /3
100
Aaron Tay @aarontay.bsky.social · 05/10/2026
My first reaction is, you should have a advanced mode, more transparency is always good! But then when I looked into it closer I started to have some empathy for their decision... What I did was to capture the network traffic using Chrome devtool and it gave me a WEALTH of info under the hood /2
110
Aaron Tay @aarontay.bsky.social · 05/10/2026
was talking to a search vendor providing new AI search mode, as usual I complained that while parts of the hybrid search was transparent - they show the lexical boolean expansion but the semantic search had not details. Intriguing he said they USED to show more but feedback was it was confusing /1
161
Aaron Tay @aarontay.bsky.social · 05/10/2026
Looks crazy promising. The actual study screened 12,293 unique records from 14 databases to get similar performance. With new openalex OQL titles abstract keyword and country you only need screen 8k!
020
Aaron Tay @aarontay.bsky.social · 05/10/2026
You are definitely better at translation than me. I'm a noob. But if you really got 8k results with 127 gold standard retrieved. Case closed really at least for this set.
130
Aaron Tay @aarontay.bsky.social · 04/10/2026
Yes. The part that is confusing is how cursor based ranking works. My guess is currently it might just rank some and then throw up its hands and just go by workid. OpenAlex doc does say lexical scoring includes proximity so this may have led to improvement in ranking vs my original constructed oql
000
Aaron Tay @aarontay.bsky.social · 04/10/2026
Complacency that leads one to go around calling themselves "superstars" or "rock stars" or cringe lines like "librarians have a super power" or "you can believe me, I am a librarian"
051
Aaron Tay @aarontay.bsky.social · 04/10/2026
Complacency that self comforts ourselves we already have all the skills and knowledge needed without learning new skills to face some new issue (what's the chances this is true?).
110
Aaron Tay @aarontay.bsky.social · 04/10/2026
What i can say is complacency is what will kill us. Complacency that we and we alone have or worse already have the answers to a novel, emerging, hard multi faceted problem. Complacency that goes because this was the way we did things for decades means this is how it should be.
131
Aaron Tay @aarontay.bsky.social · 04/10/2026
Now librarians of my generation are at the wheel and we face a even harder challenge than the Internet. I can tell you I have no clear idea what the solution is either !
110
Aaron Tay @aarontay.bsky.social · 04/10/2026
To be fair at the time i was reading some accounts of librarians at the time reflecting on the wrong turns made in 90s as Web emerged & if I were a senior person with influence at the time i wouldn't have done better. But I just instinctively dislike the sense of complacency
100
Aaron Tay @aarontay.bsky.social · 04/10/2026
Eg "Librarians will survive! People said Internet will kill us in 90s/2000s but we still here". My response was "sure we survived but going from a world where library is centre of knowledge to an after thought for most is not exactly covering yourself with glory" more carefully phrased of course
110
Aaron Tay @aarontay.bsky.social · 04/10/2026
I remember some senior who i respected said I was the only panelist that was "real". What I remember being was that i was candid and frank mostly because I was triggered a bit by overly rah rah talk particularly by one panelist (a much older and senior person)
100
Aaron Tay @aarontay.bsky.social · 04/10/2026
Recently I was asked to be on a local panel on future of librarians. This sounded familar then it clicked. In the 2012 iteration of the same conference, I was on pretty much the same panel! I even found some notes on how I intended to respond..
100
Aaron Tay @aarontay.bsky.social · 04/10/2026
Take this slide from September 2023 I presented at EIFL General Assembly where I predicted the next 3 years. I recently revisited this at EIFL 2026 Virtual.. How did i do? /1
000
Aaron Tay @aarontay.bsky.social · 04/10/2026
I'm old enough in the profession I can compare what I said or thought about things years if not decades ago...Really interesting /1
240
Aaron Tay @aarontay.bsky.social · 04/10/2026
@jevbot.bsky.social I sometimes worry that librarianship is responding to AI by defining our future as teaching AI literacy. That is too small. We shouldn't just teach people how to use systems built by somebody else.
160
Aaron Tay @aarontay.bsky.social · 04/10/2026
Will need to test on many cases also. Stansfield 2025 strikes me as a clearly favourable case for openalex with high recall and relatively low records returned even pre OQL features. /end
000
Aaron Tay @aarontay.bsky.social · 04/10/2026
Currently OQL proximity + keyword almost blows up the results so if you are of the screen 100% school you might not want that. But in this case at least it helps with recall @ orginal workload. There are plans to do crosswalk with mesh... not sure how that will change things..exciting times
100