VideoLAN Developer Days 2025 London Report Part 1. (en)
This is the 14th event hosted by VideoLAN Organizations. Named as VideoLAN Developer Days 2025 (VDD) which is held in London. Start from…
VideoLAN Developer Days 2025 London Report Part 1. (en)

Reported by Takesato Hayashi
This is the 14th event hosted by VideoLAN Organizations. Named as VideoLAN Developer Days 2025 (VDD) which is held in London. Start from the Community Day on 31st October at the same timing of Halloween. The main conference was held on 1st November and 2nd November at University College London (UCL). This is my fourth time attending VDD, and I’m really looking forward to seeing the members again, too. I traveled from Haneda, Tokyo (NHD) to Incheon Airport (ICN), Korea first, then ICN to London Heathrow (LHR). A total 17 hours 20 minutes flight.
# The Community Day? You know, open-source projects like VLC media player and FFmpeg are developed by various contributors around the globe. So, most of the communication is online. Physical event VDD is the great opportunity for everyone involved in the open-source project to gather in one place to meet each other. This pre-event provides them to know each other and work together experience throughout the day. Every time VDD is hosted in different countries and cities. So, you must use your imagination and knowledge to explore the city with your teammate.
Day 1: Community Day in London on 31st October In the morning, people arrived at Holiday Inn London — Regent’s Park at 9:30 AM. The President of VideoLAN, Jean-Baptiste Kempf (jb) announced the start of the day. Form a group of people at least 5, then receive a mission with the special map. Create a name of the group, start! If you walk around all the places, it will be more than 25 km.


jb announced the start of the day!! / Team Paddington!!
Everyone received a map, 3 pages assignment which includes General hunting + Cat’s hunting and Secret Key hunting.
We explore the city to see and pass through The Lions at Trafalgar Square, The changing of the King’s Life Guard ceremony, The Royal Courts of Justice, The London eye, The River Thames, St Paul’s Cathedral, Faryners House and Leadenhall Market. As always in London, we had rain throughout the day, so my team “Paddington” stopped our exploration at 4:30 PM. Based on my watch records, I walked 14.6 km with 20,920 steps.



We explored the streets of London with Team Paddington!
After walking a lot, everyone gathered again at the bar “BrewDog” to celebrate the great achievement of the day. When we walked back to the hotel, the town of London had a variety of people who were wearing costumes. We are spotted by three heroes and ask for the picture together. Quite an interesting thing is that hero in blue suites knows how VLC media player supports the world of multimedia 😉


A Halloween night in London 🎃💀🍷
Day 2: VideoLAN Developer Days 2025 in London at UCL on 1st November The two-day conference begins today at University College London (UCL). Registration starts at 9:00 am and jb’s opening remark starts at around 9:30 am.


jb’s Opening Remarks / Traditional VDD Group shot!!
The event commenced with expressions of gratitude to UCL, the host institution for this conference, to MUX, the sponsor, and to Vibhoothi, the venue coordinator, following the previous conference in Dublin two years ago. The morning featured five sessions, while the afternoon sessions were split into two rooms: one for VLC and another for AV2/FFmpeg. Finally, jb told everyone present at today’s venue that they are members of the VideoLAN community and, at the same time, members of the widely used Multimedia FLOSS community around the world, and to be proud of that. Every time I attend, I think VDD is a gathering of the members who support so much of the multimedia technology we use every day. It’s amazing, isn’t it?
Machine Learning and Reinforcement Learning: Distinct Trajectories, Underlying Convergence by LASP UCL A warm welcome and presentation from the Learning and Signal Processing Team (LASP) at UCL, introduced the LASP Team’s research on Human-Centric XR, Graph-Based Machine Learning, and Reinforcement Learning / Autonomous Systems.
One particularly striking question in the discussion of Immersive Reality was “Is VR sustainable?”. If you think about prioritizing the user experience, ultra-low latency communication is essential, and rendering consumes significant computational resources in the space beyond the user’s visible screen.



LASP Team
For example, a 5-minute 360-degree video (150Mbps) consumes approximately 5.5GB of data. While it’s essential to send the optimal video tailored to each user, the presentation also emphasizes the importance of minimizing the data transmitted.
libswscale reimagined by Niklas Haas (PDF) Niklas is an independent consultant, is the author of the library the rendering library “libplacebo” and contributor to projects like mpv, FFmpeg, dav1d. He explained the recent evolution of libswscale in 2024 and 2025, then looked at the future direction.
libswscale<https://ffmpeg.org/libswscale.html>

Niklas Haas
libswscale, 2024 In the first part of his presentation, Niklas Haas delivered a detailed critique of the existing library’s structural flaws. The core issue was that the entire conversion pipeline (scaling, colour space, pixel format) was tightly integrated into one massive, highly complex structure called the SwsContext. As he stated, the API was “highly stateful,” forcing developers to manually manage large amounts of interdependent settings. This architectural complexity ultimately culminated in serious quality issues like “random bugs”, “unpredictable behavior” and the inability to extend the code. The conclusion was clear, the system required some “cleanups” required.
libswscale, 2025 Second part of his presentation, starting with “Stateless public API”, the solution was a complete architectural shift away from the single, stateful context. The new design is centered around breaking down the conversion process into small, independent building blocks called SwsOp. These operations are then efficiently chained together. Also mentioning new helpers SwsGraph + SwsPass making the conversion path transparent and optimised.
Future direction Outlined the ambitious areas for future development and collaboration, focusing on performance, expansion, and community contribution. Future efforts are focused on performance optimisation through “Runtime SIMD generation” for architectures like “ARM NEON” and “RISC-V”, and developing a “GPU backend” by compiling “SwsOps → SPIR-V → Vulkan.” The roadmap includes adding “Missing functionality” such as better “Scaling, LUTs, subsampling, Bayer, EOTF/OETF, …” He called for community help via “Moar SIMD”, “Testing” (using the “SWS_UNSTABLE” flag), and “Funding”.
SVT-AV1-PSY by Julio Barba & Gianni Rosato (PDF) Gianni Rosato is a student at Worcester Polytechnic Institute(WPI) and an open-source enthusiast, while Julio Barba is a developer specializing in video and image compression technology like AV1 and AV2. Together, they started SVT-AV1-PSY, a passion project designed to optimise modern video encoding and a fork of the SVT-AV1 encoder with a focus on look as good as possible to the human eye. They presented the project’s history, its “psychovisual” feature set, and its ultimate goal of integrating these community-driven improvements back into the mainline SVT-AV1 encoder.
SVT-AV1-PSY<https://svt-av1-psy.com/>

Julio Barba & Gianni Rosato
The “Why”. Addressing the Metrics Gap. The presentation began by analyzing the historical shift in encoder development, noting a transition from the “perceptual focus” of x264 to a “metrics focus” in newer encoders like SVT-AV1. They highlighted that the community “wasn’t satisfied with AV1 encoders” because libaom, while flexible, was “very slow,” and SVT-AV1 was “benchmaxxed” to target the arithmetic mean of metrics like PSNR, SSIM, and VMAF rather than subjective visual quality. Furthermore, existing libaom forks suffered from a “messy commit history” and an “inscrutable release cycle”. To solve this, they created SVT-AV1-PSY to “bring back community encoder development”, aiming to recruit developers and implement “ergonomic features” that are easy to use and understand.
Psychovisual Features A significant portion of the talk detailed the “20+ impactful features” introduced to improve visual fidelity. A key innovation and their first feature “Variance adaptive quantization” (Variance Boost), which improves “low-contrast scenes” and helps “retain faint textures” such as skin. They also addressed granularity issues with “Extended and quarter-step CRF,” extending the range to 70.0 in “0.25 increments” to prevent file sizes from becoming too large at high CRFs. Another major addition is “AC Bias,” an “energy-preserving in-loop filter” where lower strengths help retain sharpness and higher strengths improve “film grain retention” significantly. Finally, “Tune 4 ‘Still Picture’” treats AV1 as an image codec, optimizing for “good perceptual metrics” like SSIMULACRA2 & Butteraugli by combining quantization matrix scaling and DLF sharpness controls.
Mainlining and Future Impact The final section focused on the project’s success in contributing back to the wider ecosystem, specifically the “work to merge PSY code to upstream SVT-AV1”. They noted that “collaboration and feedback” from the main SVT-AV1 team have been “outstanding” with features like “Variance Boost”, “Tune 4” and “Adaptive Film Grain Synth” already merged or planned for upstream integration. Beyond SVT-AV1, the initiative has influenced other projects; the “entire Tune IQ feature set” was ported to libaom for avifenc, and improvements to screen content detection and block hashing are helping improve “AV2 coding tools” for future still picture tuning.
Enabling Intelligent Media Playback on RISC-V — Running VLC by Yuning Liang Hong Kong based and formed in 2022, DeepComputing is a pioneer dedicated to advancing RISC-V adoption beyond existing chipsets. Led by founder Yuning Liang, the company focuses on driving the RISC-V ecosystem.

Yuning Liang
Hardware Evolution & Framework Partnership Yuning Liang showcased the rapid evolution of RISC-V hardware, moving from the 2023 DC-ROMA laptop to a new strategic partnership with Framework for modular DIY laptops.
- DC-ROMA II Mainboard (Oct 2025): Promises an 8-core 2GHz processor, support for up to 64GB RAM, and a 50 TOPS AI accelerator, targeting a price under $300.
Chiplet Solution To solve the challenge of “unknown required compute power”. He introduced the ESWIN 7702x, the world’s first RISC-V Chiplet AI SoC. It features a 2-die architecture combining CPU, GPU, and NPU to deliver 50 TOPS of AI compute and 8K encoding capabilities.
Software: Porting VLC & AI His team undertook the difficult task of porting VLC media player to RISC-V while integrating local AI features.
- Architecture: Runs VLC (media), Whisper (speech-to-text), and LLMs (translation) natively.
- Challenges: Despite handling 125+ dependencies, performance is still being optimised; the DeepSeek 7B model currently runs at a slow 5 tokens per second.
Ecosystem Call to Action DeepComputing is upstreaming its AI support frameworks and launching a sponsorship program. They aim to support “100 AI Startups” by providing free Framework devices with AI compute to active contributors.
Reconstructing 3D from Compressed Video: An AV1-Based SfM Pipeline by Julien Zouein Julien Zouein is a PhD Student and Research Assistant at Trinity College Dublin under Prof. Anil Kokaram, and a co-founder of Kyber. With a background in real-time video interaction and Cloud Gaming, his research focuses on optimizing Real-Time Video Processing by reusing available features from encoded bitstreams. He is also a member of the VideoLAN Organization.


Julien Zouein
The Concept: AV1 as a Feature Matching Tool Standard Structure from Motion (SfM) pipelines usually require expensive image processing on raw pixels to reconstruct 3D scenes. Julien presented an approach that transforms the AV1 encoder itself into a feature matching tool for high-quality SfM. The core innovation lies in working directly on the compressed stream without the need for traditional image processing. By extracting features such as Motion Vectors, Reference Frames, and Block Sizes directly from the AV1 bitstream, the pipeline bypasses the most computationally heavy steps of 3D reconstruction.
Pipeline The proposed pipeline operates by extracting motion data and partitioning info from the AV1 bitstream to generate a motion field. This data is then used to perform correspondence searches and keypoint matching, feeding into an incremental reconstruction module.
Performance and Results The results of his method are significant for both efficiency and quality. By leveraging the compressed domain.
- Efficiency: CPU utilization is reduced by 95% and overall processing time is reduced by 42% compared to traditional methods.
- Quality: Technique successfully generates a denser 3D point cloud, boasting 8x more 3D points than standard sparse reconstruction methods.
- Visual Fidelity: Comparisons against standard methods (like SIFT + Exhaustive Matching) show that the AV1-based approach captures significantly more detail in the environment, such as building textures.
Impact and Future Release This optimization unlocks new applications for low-power edge devices, such as drones, where CPU constraints previously limited onboard 3D calculations. He plans a public release of the optimised code later this year, associated with this paper “Leveraging AV1 motion vectors for Fast and Dense Feature Matching”.
# Afternoon Conference In the afternoon, we were split into two groups. Development meetings for VLC media player, Open-Codecs (SVT-AV1-PSY, AV2) + FFmpeg. I attended both Open-Codecs and the VLC Media Player meeting.
Serious discussions were held regarding each future topic. Participants came from diverse backgrounds, and it was both fascinating and unique to see multiple perspectives emerge one after another regarding a single proposal or topic.
# Community Dinner After both meetings that went on until after 6 pm, the first night of the conference was the much-anticipated VideoLAN community dinner.



Community Dinner at The Marquis Cornwallis
Held at The Marquis Cornwallis<https://www.themarquiscornwalliswc1.co.uk/>, a British pub serving traditional food and drinks, the dinner was the perfect place to gather and celebrate the results of yesterday’s Community Day. The winning team was presented with LEGO as a prize. Congratulations!!
To be continued in Part 2<https://medium.com/tokyo-video-tech/videolan-developer-days-2025-london-report-part-2-en-7818cdde9c66>
메타데이터
- post_id
- f161475e7331
- slug
- videolan-developer-days-2025-london-report-part-1-en-f161475e7331
- url
- https://medium.com/tokyo-video-tech/videolan-developer-days-2025-london-report-part-1-en-f161475e7331
- canonical_url
- https://medium.com/tokyo-video-tech/videolan-developer-days-2025-london-report-part-1-en-f161475e7331
- author_url
- https://medium.com/@sabaneko
- status
- ok
- fetched_at
- 2026-07-26 23:43:52