Sign in

Ashvanth.S

@ashvanths.bsky.social
70 followers 84 following 61 posts

Deep Learning Practitioner | Language Lead for Tamil @ HuggingFace | Interested in Continual Learning and Generative Models | Website : ash-01xor.github.io X : twitter.com/ashvanth_s1

PostsRepliesMedia
Ashvanth.S @ashvanths.bsky.social · 26/01/2026
"Code appears. It looks right. The diff gets approved. But no one held it in their head. No one walked through the house; they just glanced at a photograph and called it home." Wrote about coding, Sisyphus, and building where the line keeps moving. link: asherr.bearblog.dev/in-pursuit-of/
asherr.bearblog.dev
In pursuit of
In the span of ten seconds, a fleeting moment, lives can be transformed forever. 100 Meters (Hyakuemu), produced by Rock 'n' Roll Mountain, is a film based u...
010
Ashvanth.S @ashvanths.bsky.social · 11/05/2025
Have you ever felt like you lost your focus while reading a book and wandered into deep internet rabbit holes? Introducing sollu : AI-powered dictionary. Uses the Gemini model under the hood. It is open-sourced as well :).
000
Ashvanth.S @ashvanths.bsky.social · 26/02/2025
Quite a humbling experience every day while coding. You start with an issue and a vision about how to solve the problem and then pretty much the road traveled often to reach the solution isn't straightforward. Humbled each and every day to understand and accept and that it is how it is.
000
Ashvanth.S @ashvanths.bsky.social · 26/02/2025
Pretty similar to how Jio first gained share of the internet users in India. Interesting to note big companies have the ability to shell out too much to develop and operate to gain market share. Only time shall tell what this will lead to
010
Ashvanth.S @ashvanths.bsky.social · 19/02/2025
Ahh finally a blog post from you , it is quite difficult to maintain a site right like publishing frequent posts
020
Ashvanth.S @ashvanths.bsky.social · 19/02/2025
Over a period of time , getting to realize that im having my flow states during certain periods of time and getting to schedule tasks around it. Guess the goal is to build systems that can make sure we enter such states like on and off button.
000
Ashvanth.S @ashvanths.bsky.social · 13/02/2025
Not able to point of the difference particularly , but gpt-4o-mini seems to work way too fast over the last day. From taking around 4 to 5 mins to process a 65-page PDF for extraction, it takes around 3 mins. Do you guys want me to run benchmark tests and probably write a blog post about it ?
010
Ashvanth.S @ashvanths.bsky.social · 13/02/2025
Looking forward to the next unit of the Agents course and building more @benburtenshaw.bsky.social @hf.co
020
Reposted by Ashvanth.S
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/02/2025
I just finished writing up my take on reasoning models: magazine.sebastianraschka.com/p/understand... Here, I 1. Discuss the advantages & disadvantages of reasoning models 2. Of course, describe and discuss DeepSeek R1 3. Describe the 4 main ways to building & improving reasoning models
magazine.sebastianraschka.com
Understanding Reasoning LLMs
Methods and Strategies for Building and Refining Reasoning Models
39020
Ashvanth.S @ashvanths.bsky.social · 03/02/2025
Slowly building it one at a time. Thanks to @sebastianraschka.com for his book. Implementing things from scratch takes a lot of time , but valuable experience. github.com/ash-01xor/Re...
github.com
GitHub - ash-01xor/Rebuild-LLM: Building Large language model from scratch
Building Large language model from scratch. Contribute to ash-01xor/Rebuild-LLM development by creating an account on GitHub.
010
Ashvanth.S @ashvanths.bsky.social · 03/02/2025
Building SmolGPT myself , have plans to extend it. but before that struggling with managing python versions !!! Had to use pyenv and then pip. like now i get why experienced devs are frustrated with python package management
000
Ashvanth.S @ashvanths.bsky.social · 02/02/2025
Updated my site after quite a long time also added a note for how to update your arch linux system. Do check it out if you use arch or if you like to as well :)
010
Ashvanth.S @ashvanths.bsky.social · 07/01/2025
Only a few more annotations are needed to complete the initial goal. for Tamil. Do join the initiative alongside me , your contribution is highly valuble
000
Ashvanth.S @ashvanths.bsky.social · 07/01/2025
Interesting to see the hype of agents and using them , but almost everyone who uses the term throws it away just like that. All I get to see is a clearly well-defined workflow in a constrained environments most of the time and yet they are being called 'agents'.
000
Ashvanth.S @ashvanths.bsky.social · 05/01/2025
Got to find this today only in Python when i made a typo by mistake. How does the for loop work when i present the number inside range like that ??
000
Ashvanth.S @ashvanths.bsky.social · 01/01/2025
Happy new year sebastian !! was waiting for the post
010
Ashvanth.S @ashvanths.bsky.social · 30/12/2024
Well, we are halfway through our initial goal of the Fineweb-C sprint for Tamil. Hopefully I would love to complete the initial goal of annotating 1000 texts within the next two days Do join if you would like to contribute! data-is-better-together-fineweb-c.hf.space/share-your-p...
data-is-better-together-fineweb-c.hf.space
tam - தமிழ் - Tamil
Join and contribute to the dataset tam - தமிழ் - Tamil
000
Ashvanth.S @ashvanths.bsky.social · 23/12/2024
Since being used to python development from the start i dont think i never had an issue using pyenv , venv , conda etc. Like it never felt like a chore. But then hearing about devs from other communities really does make me question why .
010
Ashvanth.S @ashvanths.bsky.social · 20/12/2024
got to read that alec radford left open ai , like what is even happening at open ai
020
Ashvanth.S @ashvanths.bsky.social · 17/12/2024
which in itself is based on the success of their previous films. As risks taken decreases due to a formulaic process , so does the excitement and the curiosity.
000
Ashvanth.S @ashvanths.bsky.social · 17/12/2024
the big names present in the resume is overlooked as a factor of judgement for their talent or in making films where rather than the concept or story , the focus shifts to the kind of artists brought in to play the characters , their star power and influence to bring audience to theaters ...
100
Ashvanth.S @ashvanths.bsky.social · 17/12/2024
Somehow deep down i always get to think about how optimization of any process leads to boredom over a period of time. The excitement and the risks once taken might decreases due to the numbers the clouds our judgement. Like while recruiting , where folks are given standard questions to solve or ..
100
Ashvanth.S @ashvanths.bsky.social · 14/12/2024
Big thanks to @dvilasuero.hf.co , @nataliaelv.hf.co and team 🙌. Would love to see more people join this effort
010
Ashvanth.S @ashvanths.bsky.social · 14/12/2024
Well, around 10 percent of the initial goal is complete, and so far, it's been quite a one-man army effort. We're still in the hunt for more people to join and contribute to this open-source initiative. @hf.co data-is-better-together-fineweb-c.hf.space/share-your-p...
data-is-better-together-fineweb-c.hf.space
tam - தமிழ் - Tamil
Join and contribute to the dataset tam - தமிழ் - Tamil
141
Ashvanth.S @ashvanths.bsky.social · 13/12/2024
The process has just begun, and we are actively seeking collaborators for Tamil. Join us in this open-source initiative! Building better models demands a better annotation process, and we are deeply committed to achieving this together data-is-better-together-fineweb-c.hf.space/share-your-p...
data-is-better-together-fineweb-c.hf.space
tam - தமிழ் - Tamil
Join and contribute to the dataset tam - தமிழ் - Tamil
010
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
While these are like the summary of what he considers to be the trends going on right now , interesting to note how it might span out in the future. Looking forward to building now !
010
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
- Enterprise Search: Integrating LLMs with search capabilities empowers intelligent assistants to manage vast knowledge bases effectively. - Assistant Applications: These solutions improve workflows by providing accurate, context-aware information.
100
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
- Support Customization: Adaptation of models to domain-specific data for optimal performance. Trend 3: The Convergence of LLMs and Search Large language models (LLMs) and search are increasingly intertwined, revolutionizing information retrieval: ....
100
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
Trend 2 : Platform Choice matters The right platform can determine the success of AI initiatives. Enterprises benefit from platforms that: - Provide Pretrained Models : Easy access to SOTA models. - Enable Production Management: Seamless monitoring and scaling in real-world deployments....
100
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
Democratization: AI tools are increasingly accessible, enabling more people to develop AI without extensive resources. Generalization Across Tasks: The shift towards universal models capable of performing millions of tasks replaces the need for task-specific models....
100
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
Trend 1 : AI is Accelerating** AI development is speeding up, and breaking barriers in scale and accessibility. Key advancements include: - Data Efficiency: Models require less training data thanks to improved algorithms and pretraining paradigms that leverage foundational knowledge....
100
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
Its been great few weeks reading about agents after going through the course conducted by Berkeley. While there was lots of insightful talks , one that was particularly insightful was on week 4 , where Burak Gokturk got to talk about enterprise trends about agents. The most important trend...
100
Ashvanth.S @ashvanths.bsky.social · 12/12/2024
coding when someone is watching is quite a nervous experience. all of sudden there is quite a bit of fumbling , struggling to come up with names..
000
Ashvanth.S @ashvanths.bsky.social · 11/12/2024
This would not have been possible without the joint vision of our team. Thanks to each one of them for contributing.
010
Ashvanth.S @ashvanths.bsky.social · 11/12/2024
An open-source effort at its fullest: open weights, open data, open code. Hugging Face: huggingface.co/maya-multimo... Code: github.com/nahidalam/ma... Thanks to Cohere for AI as well for supporting us.
github.com
120
Ashvanth.S @ashvanths.bsky.social · 11/12/2024
Introducing Maya – A New Multimodal Multilingual Vision-Language Model. Maya is a completely open source, open weight, and open dataset, designed to handle 8 languages, cultural diversity, and nuanced real-world contexts in vision-language models. Paper: arxiv.org/abs/2412.07112
arxiv.org
Maya: An Instruction Finetuned Multilingual Multimodal Model
The rapid development of large Vision-Language Models (VLMs) has led to impressive results on academic benchmarks, primarily in widely spoken languages. However, significant gaps remain in the ability...
130
Ashvanth.S @ashvanths.bsky.social · 10/12/2024
Vanakkam makkalae , glad that I’ll be leading the FineWeb 2 collaborative annotation sprint for Tamil! 🤗 I’ll be helping to build an open dataset to improve language models for our language. Do join the process of improving models ! huggingface.co/spaces/Huggi... huggingface.co/spaces/data-...
021
Ashvanth.S @ashvanths.bsky.social · 09/12/2024
Exciting things coming up @hf.co . Can't wait to reveal tomorrow.
010
Ashvanth.S @ashvanths.bsky.social · 09/12/2024
For most lexical based engines , it just works and if your dataset size is lower even more reasons to pursue it. Yes using S-BERT and other fancy methods might seem like good idea but the process is to ensure that you have simple baseline first and then all these fancy methods come up.
000
Ashvanth.S @ashvanths.bsky.social · 09/12/2024
Weekend was quite cool , got to just keep my head down and work on retrieval engines over a custom database. Implemented thing right from TF-IDF to RAG and saw quite some interesting things. Anyone who says TF-IDF is not the best must need to implement it and check it first...
100
Ashvanth.S @ashvanths.bsky.social · 07/12/2024
Everyday as i code , i find more hidden details and my shortcomings.
000
Ashvanth.S @ashvanths.bsky.social · 05/12/2024
Congratss Lucas , love how you are doing this FAQ both on X as well as here :)
010
Ashvanth.S @ashvanths.bsky.social · 04/12/2024
4th day and im already feeling the difficult of it rising. "Ceres Search" - Day 4 - Advent of Code 2024 #AdventOfCode adventofcode.com/2024/day/4
adventofcode.com
Day 4 - Advent of Code 2024
030
Ashvanth.S @ashvanths.bsky.social · 04/12/2024
Yes indeed , made a PR already 😀.
000
Ashvanth.S @ashvanths.bsky.social · 04/12/2024
It is more about understanding the process used to train the model using TRL and other libraries potentially. Might also write a blog on it , do let me know if you are interested to read on it (might help me get started to work on it :) )
010
Ashvanth.S @ashvanths.bsky.social · 04/12/2024
Cool work done by huggingface @benburtenshaw.bsky.social . Loving the smol-course , got to train a small model on the stack dataset (just on the python subset of it) and upload it to the hub. Link to the model: huggingface.co/collections/... The idea is not to use the model trained right away ....
huggingface.co
Smol-Course-Models - a peaceAsh Collection
The collection of models trained for different chapters and datasets provided in the smol-course
210
Ashvanth.S @ashvanths.bsky.social · 03/12/2024
Got to complete "Mull It Over" - Day 3 - Advent of Code 2024 #AdventOfCode adventofcode.com/2024/day/3
adventofcode.com
Day 3 - Advent of Code 2024
010
Ashvanth.S @ashvanths.bsky.social · 02/12/2024
Yep got to complete "Red-Nosed Reports" - Day 2 - Advent of Code 2024 #AdventOfCode adventofcode.com/2024/day/2
adventofcode.com
Day 2 - Advent of Code 2024
010
Ashvanth.S @ashvanths.bsky.social · 02/12/2024
Hiiii Zach , hope you are doing fine !
000
Ashvanth.S @ashvanths.bsky.social · 01/12/2024
Lol saw the leaderboard , someone solved it within 9 seconds . Like someone have an entire pipeline setup and read to download data , send it to LLM and submit it back. Damn quite a surprise.
010